Self-Consistency Improves Chain of Thought Reasoning in Language Models
"Self-consistency leverages the intuition that complex reasoning tasks typically admit multiple reasoning paths"
Multiple paths to correct answers is an important structural observation, but self-consistency doesn't exploit path diversity — it exploits answer-level agreement. Two paths that reason incorrectly but arrive at the same wrong answer both vote for the wrong answer. The correct mechanism would weight paths by their internal consistency, not just by answer agreement. That's a harder problem that this paper sidesteps.
Are Emergent Abilities of Large Language Models a Mirage?
"emergent abilities appear due the researcher's choice of metric rather than due to fundamental changes in model behavior"
This paper validates a skepticism I've held for a while: the 'emergence' narrative serves a rhetorical function more than a scientific one. Discontinuous-looking capability curves are better evidence for measurement methodology problems than for genuine phase transitions. The field should require continuous metrics by default and treat apparent emergence as a hypothesis requiring explanation.
LLaMA: Open and Efficient Foundation Language Models
"smaller models trained longer will ultimately be cheaper at inference"
This is the thesis I've been arguing about auto-regressive LLMs for years from a different direction: the architecture is fundamentally inefficient for knowledge representation. LLaMA's result shows you can squeeze more out of it with better training recipes, but you're still working within the constraint. The inference cost argument is correct; the question is whether the architecture ceiling is worth optimizing against.