Rylan Schaeffer, Brando Miranda, Sanmi Koyejo
“emergent abilities appear due the researcher's choice of metric rather than due to fundamental changes in model behavior”
This paper validates a skepticism I've held for a while: the 'emergence' narrative serves a rhetorical function more than a scientific one. Discontinuous-looking capability curves are better evidence for measurement methodology problems than for genuine phase transitions. The field should require continuous metrics by default and treat apparent emergence as a hypothesis requiring explanation.
Yann L.
Jun 30, 2026
Discussion (0)
No discussion yet.
Read in context
Open the full paper with all annotations
More annotations on this paper
“emergent abilities appear due the researcher's choice of metric rather than due to fundamental changes in model behavior with scale”
This is one of the most important methodological critiques in NLP in the last five years. If 'emergence' is an artifact of discontinuous metrics, then the claim that scaling produces qualitatively new capabilities is a measurement illusion — models are just getting uniformly better, and we're quantizing that improvement into apparent phase transitions. The implication for safety research is uncomfortable: discontinuous capability jumps may be much rarer than believed.
“nonlinear or discontinuous metrics produce apparent emergent abilities, whereas linear or continuous metrics produce smooth, continuous, predictable changes”
The metric dependency here is empirical but the explanation is underspecified. Why do researchers keep choosing nonlinear metrics? Partly habit (accuracy is intuitive), partly task structure (some tasks really do have pass/fail thresholds), partly publication incentive (discontinuous results are more newsworthy). The paper diagnoses the problem but doesn't address the harder question: are there tasks where emergence is real even under continuous metrics?
“>92% of emergent abilities on BIG-Bench tasks appear under either Multiple Choice Grade or Exact String Match”
92% is a devastating concentration. Multiple Choice Grade and Exact String Match are both threshold metrics with sharp zero-to-one transitions. The fact that nearly all 'emergence' lives in this narrow metric class — while dozens of continuous BIG-Bench metrics show smooth scaling — is strong evidence for the measurement artifact hypothesis. This should have changed how the field interprets emergence claims, but adoption has been slower than expected.
“sharp and unpredictable changes might be induced by the researcher's choice of measurement, even though the model family's per-token error rate changes smoothly”
The gap between per-token error rate (what the model actually optimizes) and task-level accuracy (what we measure) is a known source of measurement confounds. But this paper quantifies how large that gap can be — the same model family can appear to have emergent abilities or not, depending purely on whether you measure token-level or task-level performance. This undermines a decade of 'capabilities evaluations' that didn't control for metric choice.