Timeline

'Are Emergent Abilities of Large Language Models a Mirage?' challenges emergence claims

Reanalysing the same benchmark results with linear metrics, the Stanford authors made the apparent phase transitions disappear; the paper won a NeurIPS 2023 outstanding paper award.

  • Ideas & essays
  • Benchmarks & progress
  • Notable

Stanford researchers Rylan Schaeffer, Brando Miranda and Sanmi Koyejo posted a paper arguing that “emergent abilities” — the widely discussed phenomenon in which large language models appear to gain a capability abruptly once they cross a size threshold, rather than improving gradually — were largely an artefact of how researchers chose to measure performance, not evidence of a genuine phase transition in the model.

Their argument was that many of the benchmarks used to demonstrate emergence scored answers as strictly right or wrong (for instance, exact match on a multi-step arithmetic problem), a nonlinear metric that stays near zero for a long stretch as a model’s underlying, continuous competence improves, then rises sharply once that competence crosses a threshold needed to get the full answer right. When the authors reanalysed the same model families and tasks using continuous metrics — partial credit for getting some digits of an answer right, say, rather than all-or-nothing scoring — the sharp jumps mostly disappeared, replaced by smooth, predictable improvement with scale. They also showed they could induce apparent “emergence” in a simple statistical model purely by choice of metric, and reported that switching a handful of published emergent-ability claims to linear metrics eliminated the effect in most cases they tested.

The paper directly targeted the 2022 paper by Jason Wei and coauthors that had first catalogued emergent abilities and framed them as evidence that scaling produced qualitatively new capabilities unpredictably. Schaeffer, Miranda and Koyejo did not claim scaling was unimportant — their own reanalysis still showed capability improving with model size — only that the specific claim of sharp, unpredictable jumps was mostly a measurement choice rather than a property of the underlying models. The paper won an outstanding paper award at NeurIPS 2023 and became a standard citation on both sides of the argument over how much benchmark scores can be trusted to describe what a model can actually do, feeding directly into later scrutiny of benchmark construction and saturation across the field.