Are Emergent Abilities of Large Language Models a Mirage?

Rylan Schaeffer, Brando Miranda, Sanmi Koyejo

Introduction

Emergent properties of complex systems have long been studied across disciplines, from physics to biology to mathematics. The idea of emergence was popularized by Nobel Prize-winning physicist P.W. Anderson’s “More Is Different” (anderson1972more, ), which argues that as the complexity of a system increases, new properties may materialize that cannot be predicted even from a precise quantitative understanding of the system’s microscopic details. Recently, the idea of emergence gained significant attention in machine learning due to observations that large language models (LLMs) such as GPT (brown2020language, ), PaLM (chowdhery2022palm, ) and LaMDA thoppilan2022lamda exhibit so-called “emergent abilities” (wei2022emergent, ; ganguli2022predictability, ; srivastava2022beyond, ; brown2020language, ) (Fig. 1).

The term “emergent abilities of LLMs” was recently and crisply defined as “abilities that are not present in smaller-scale models but are present in large-scale models; thus they cannot be predicted by simply extrapolating the performance improvements on smaller-scale models” wei2022emergent . Such emergent abilities were first discovered in the GPT-3 family brown2020language . Subsequent work emphasized the discovery, writing that “[although model] performance is predictable at a general level, performance on a specific task can sometimes emerge quite unpredictably and abruptly at scale” ganguli2022predictability . These quotations collectively identify the two defining properties of emergent abilities in LLMs:

Sharpness, transitioning seemingly instantaneously from not present to present

Unpredictability, transitioning at seemingly unforeseeable model scales

These emergent abilities have garnered significant interest, raising questions such as: What controls which abilities will emerge? What controls when abilities will emerge? How can we make desirable abilities emerge faster, and ensure undesirable abilities never emerge? These questions are especially pertinent to AI safety and alignment, as emergent abilities forewarn that larger models might one day, without warning, acquire undesired mastery over dangerous capabilities steinhardt2022future ; hendrycks2022emergent ; krakovna2022sharp1 ; krakovna2022sharp2 .

In this paper, we call into question the claim that LLMs possess emergent abilities, by which we specifically mean sharp and unpredictable changes in model outputs as a function of model scale on specific tasks. Our doubt stems from the observation that emergent abilities seem to appear only under metrics that nonlinearly or discontinuously scale any model’s per-token error rate. For instance, as we later show, >92%>92\% of emergent abilities on BIG-Bench tasks srivastava2022beyond (hand-annotated by wei2022bigbench ) appear under either of these two metrics:

This raises the possibility of an alternative explanation for the origin of LLMs’ emergent abilities: sharp and unpredictable changes might be induced by the researcher’s choice of measurement, even though the model family’s per-token error rate changes smoothly, continuously and predictably with increasing scale. Specifically, our alternative posits that emergent abilities are a mirage caused primarily by the researcher choosing a metric that nonlinearly or discontinuously deforms per-token error rates, and secondarily by possessing too few test data to accurately estimate the performance of smaller models, thereby causing smaller models to appear wholly unable to perform the task.

To communicate our alternative explanation, we present it as a simple mathematical model and demonstrate how it quantitatively reproduces the evidence offered in support of emergent abilities of LLMs. We then test our alternative explanation in three complementary ways:

We make, test and confirm three predictions based on our alternative hypotheses using the InstructGPT lowe2022instruct / GPT-3 brown2020language model family.

We meta-analyze published benchmarks srivastava2022beyond ; wei2022emergent to reveal that emergent abilities only appear for specific metrics, not for model families on particular tasks, and that changing the metric causes the emergence phenomenon to evaporate.

We induce never-before-seen, seemingly emergent abilities in multiple architectures across various vision tasks by intentionally changing the metrics used for evaluation.

Alternative Explanation for Emergent Abilities

How might smooth, continuous, predictable changes in model family performance appear sharp and unpredictable? The answer is that the researcher’s choice of a nonlinear or discontinuous metric can distort the model family’s performance to appear sharp and unpredictable.

To expound, suppose that within a model family, the test loss falls smoothly, continuously and predictably with the number of model parameters. One reason to believe this is the phenomenon known as neural scaling laws: empirical observations that deep networks exhibit power law scaling in the test loss as a function of training dataset size, number of parameters or compute (hestness2017deep, ; rosenfeld2019constructive, ; henighan2020scaling, ; kaplan2020scaling, ; gordon2021data, ; hernandez2021scaling, ; jones2021scaling, ; zhai2022scaling, ; hoffmann2022training, ; clark2022unified, ; neumann2022scaling, ). For concreteness, suppose we have a model family of different numbers of parameters N>0N>0 and assume that each model’s per-token cross entropy falls as a power law with the number of parameters NN for constants c>0,α<0c>0,\alpha<0 (Fig. 2A):

To be clear, we do not require this particular functional form to hold; rather, we use it for illustrative purposes. Let VV denote the set of possible tokens, p∈Δ∣V∣−1p\in\Delta^{|V|-1} denote the true but unknown probability distribution, and p^N∈Δ∣V∣−1\hat{p}_{N}\in\Delta^{|V|-1} denote the NN-parameter model’s predicted probability distribution. The per-token cross entropy as a function of number of parameters NN is:

In practice, pp is unknown, so we substitute a one-hot distribution of the observed token v∗v^{*}:

A model with NN parameters then has a per-token probability of selecting the correct token (Fig. 2B):

Suppose the researcher then chooses a metric that requires selecting LL tokens correctly. For example, our task might be LL-digit integer addition, and a model’s output is scored 11 if all LL output digits exactly match all target digits with no additions, deletions or substitutions, otherwise. If the probability each token is correct is independent111While the independence assumption is not true, the approximation yields results qualitatively matching the observed emergence claims., the probability of scoring 11 is:

This choice of metric nonlinearly scales performance with increasing token sequence length. When plotting performance on a linear-log plot, one sees a sharp, unpredictable emergent ability on longer sequences (Fig. 2C) that closely matches claimed emergent abilities (inset). What happens if the researcher switches from a nonlinear metric like Accuracy, under which the per-token error rate scales geometrically in target length (App. A.3), to an approximately linear metric like Token Edit Distance, under which the per-token error rate scales quasi-linearly in target length (App. A.2)?

The linear metric reveals smooth, continuous, predictable changes in model performance (Fig. 2E). Similarly, if the researcher uses a discontinuous metric like Multiple Choice Grade, the researcher can find emergent abilities (Fig. 2D), but switching to a continuous metric like Brier Score removes the emergent ability (Fig. 2F). In summary, sharp and unpredictable changes with increasing scale can be fully explained by three interpretable factors: (1) the researcher choosing a metric that nonlinearly or discontinuously scales the per-token error rate, (2) having insufficient resolution to estimate model performance in the smaller parameter regime, with resolution222Resolution is defined as “The smallest interval measurable by a scientific instrument; the resolving power.” set by 1/test dataset size1/\text{test dataset size}, and (3) insufficiently sampling the larger parameter regime.

Analyzing InstructGPT/GPT-3’s Emergent Arithmetic Abilities

Previous papers prominently claimed the GPT brown2020language ; lowe2022instruct family333As of 2023-03-15, 4 models with 350M, 1.3B, 6.7B, 175B parameters are available via the OpenAI API. displays emergent abilities at integer arithmetic tasks ganguli2022predictability ; srivastava2022beyond ; wei2022emergent (Fig. 2E). We chose these tasks as they were prominently presented brown2020language ; ganguli2022predictability ; srivastava2022beyond ; wei2022emergent , and we focused on the GPT family due to it being publicly queryable. As explained mathematically and visually in Sec. 2, our alternative explanation makes three predictions:

Changing the metric from a nonlinear or discontinuous metric (Fig. 2CD) to a linear or continuous metric (Fig. 2 EF) should reveal smooth, continuous, predictable performance improvement with model scale.

For nonlinear metrics, increasing the resolution of measured model performance by increasing the test dataset size should reveal smooth, continuous, predictable model improvements commensurate with the predictable nonlinear effect of the chosen metric.

Regardless of metric, increasing the target string length should predictably affect the model’s performance as a function of the length-1 target performance: approximately geometrically for accuracy and approximately quasilinearly for token edit distance.

To test these predictions, we collected outputs from the InstructGPT/GPT-3 family on two tasks: 2-shot multiplication between two 2-digit integers and 2-shot addition between two 4-digit integers.

On both arithmetic tasks, the GPT family displays emergent abilities if the target has 4 or 5 digits and if the metric is Accuracy (Fig. 3, top) brown2020language ; ganguli2022predictability ; wei2022emergent . However, if one changes from nonlinear Accuracy to linear Token Edit Distance while keeping the models’ outputs fixed, the family’s performance smoothly, continuously and predictably improves with increasing scale (Fig. 3, bottom). This confirms our first prediction and supports our alternative explanation that the source of emergent abilities is the researcher’s choice of metric, not changes in the model family’s outputs. We also observe that under Token Edit Distance, increasing the length of the target string from 1 to 5 predictably decreases the family’s performance in an approximately quasilinear manner, confirming the first half of our third prediction.

Prediction: Emergent Abilities Disappear With Better Statistics

We next tested our second prediction: that even on nonlinear metrics such as accuracy, smaller models do not have zero accuracy, but rather have non-zero above-chance accuracy commensurate with choosing to use accuracy as the metric. In order to accurately measure models’ accuracy, we increased the resolution by generating additional test data, and found that on both arithmetic tasks, all models in the InstructGPT/GPT-3 family achieve above-chance accuracy (Fig. 4). This confirms our second prediction. We also observe that as the target string length increases, the accuracy falls approximately geometrically with the length of the target string, confirming the second half of our third prediction. These results additionally demonstrate that the researcher’s choice of metric has the effect that one should predict accuracy to have, i.e., geometric decay with the target length.

Meta-Analysis of Claimed Emergent Abilities

Analyzing the GPT family is possible because the models are publicly queryable. However, other model families claimed to exhibit emergent abilities are not publicly queryable, nor are their generated outputs publicly available, meaning we are limited to analyzing the published results themselves ganguli2022predictability ; wei2022emergent ; wei2022bigbench . Our alternative explanation makes two predictions.

At the “population level” of Task-Metric-Model Family triplets, emergent abilities should appear predominantly on specific metrics, not task-model family pairs, and specifically with nonlinear and/or discontinuous metrics.

On individual Task-Metric-Model Family triplets that display an emergent ability, changing the metric to a linear and/or continuous metric should remove the emergent ability.

To test these predictions, we used to claimed emergent abilities on BIG-Bench srivastava2022beyond ; wei2022emergent due to the benchmark being pertinent and publicly available.

We found that most metrics used in BIG-Bench have zero task-model family pairs that exhibit emergent abilities: of the 39 preferred metrics in BIG-Bench, at most 5 display emergence (Fig. 5A). Many of the 5 are nonlinear and/or discontinuous, e.g., Exact String Match, Multiple Choice Grade, ROUGE-L-Sum (App. A.4). Notably, because BIG-Bench often scores models on tasks using multiple metrics, the lack of emergent abilities under other metrics suggests that emergent abilities do not appear when model outputs are scored using other metrics.

Because emergence score only suggests emergence, we also analyzed hand-annotated task-metric-model family triplets wei2022bigbench , which revealed emergent abilities appear with 4/394/39 metrics (Fig. 5B), and 2 metrics account for >92%>92\% of claimed emergent abilities (Fig. 5C): Multiple Choice Grade and Exact String Match. Multiple Choice Grade is discontinuous, and Exact String Match is nonlinear.

Prediction: Changing Metric Removes Emergent Abilities

To test our second prediction, we focused on the LaMDA family thoppilan2022lamda because its outputs are available through BIG-Bench. For our analysis, we identified tasks on which LaMDA displays emergent abilities with Multiple Choice Grade, then asked whether LaMDA still displays emergent abilities on the same tasks with a different BIG-Bench metric: Brier Score brier1950verification . Brier Score is a strictly proper scoring rule for predictions of mutually exclusive outcomes; for a binary outcome, the Brier Score simplifies to the mean squared error between the outcome and its predicted probability mass. LaMDA’s emergent abilities on the discontinuous Multiple Choice Grade disappeared when we changed the metric to the continuous Brier Score (Fig. 6). These results support our alternative explanation that emergent abilities are induced by the chosen metric.

Inducing Emergent Abilities in Networks on Vision Tasks

To demonstrate how emergent abilities can be induced by the researcher’s choice of metric, we show how to produce emergent abilities in deep networks of various architectures: fully connected, convolutional, self-attentional. We focus on vision tasks because abrupt transitions in vision models’ capabilities have not been observed to the best of our knowledge; this is one reason why emergence in large language models is considered so interesting. For the convolutional example, see App. B.

We first induce an emergent ability to reconstruct images in shallow (i.e., single hidden layer) nonlinear autoencoders trained on CIFAR100 natural images krizhevsky09learningmultiple . To emphasize that the sharpness of the metric is responsible for emergent abilities, and to show that sharpness extends to metrics beyond Accuracy, we intentionally define a discontinuous metric that measures a network’s ability to reconstruct a dataset as the average number of test data with squared reconstruction error below threshold cc:

Emergent Classification of Omniglot Characters by Autoregressive Transformers

We next induce emergent abilities in Transformers vaswani2017attention trained to autoregressively classify Omniglot handwritten characters lake2015human , in a setup inspired by recent work chan2022data : Omniglot images are embedded by convolutional layers, then sequences of embedded image-image class label pairs are fed into decoder-only transformers. We measure image classification performance on sequences of length L∈L\in, again via subset accuracy: 11 if all LL images are classified correctly (Fig. 8B), 0 otherwise. Causal transformers display a seemingly emergent ability to correctly classify Omniglot handwritten characters (Fig. 8C) that qualitatively matches published emergent abilities (Fig. 8A).

Related Work

Srivastava et al. srivastava2022beyond observed that while accuracy at a particular task can empirically appear sharp and unpredictable, cross entropy does not; the authors then hypothesized that emergent abilities may be partially attributed to the metric. Our paper converts their discussion into precise predictions, then quantitatively tests the predictions to reveal that: metric choice is likely wholly responsible for emergent abilities; well-known and widely-used metrics (including ones already used by srivastava2022beyond ) capture graded improvements; emergent abilities do not appear only for tasks involving multiple steps, and indeed appear most commonly on the discontinuous Multiple Choice Grade; metric choice can be used to induce emergent abilities in a novel domain (vision) in diverse architectures and tasks.

Caballero et al. caballero2022broken explain emergence by assuming a piece-wise power law functional form; under this view, emergent abilities are real, caused by a change in the governing power law. In contrast, our work suggests that emergent abilities are induced by the researcher, even under a single power law. Michaud et al. michaud2023quantization posit that emergent abilities may be real under strong data assumptions.

Discussion

Our paper presents an alternative explanation for claimed emergent abilities of large language models. For a fixed task and a fixed model family, the researcher can choose a metric to create an emergent ability or choose a metric to ablate an emergent ability. Ergo, emergent abilities may be creations of the researcher’s choices, not a fundamental property of the model family on the specific task. We emphasize that nothing in this paper should be interpreted as claiming that large language models cannot display emergent abilities; rather, our message is that previously claimed emergent abilities in brown2020language ; ganguli2022predictability ; srivastava2022beyond ; wei2022emergent might likely be a mirage induced by researcher analyses.

Our paper has several implications. Firstly, a task and a metric are distinct and meaningful choices when constructing a benchmark. Secondly, when choosing metric(s), one should consider the metric’s effect on the per-token error rate and adapt their measuring process accordingly, e.g., if one chooses accuracy, one should make sure to have sufficient data to accurately measure accuracy to avoid the risk of drawing invalid scientific conclusions. Thirdly, when making claims about capabilities of large models, including proper controls is critical. In this particular setting, emergent abilities claims are possibly infected by a failure to control for multiple comparisons. In BIG-Bench alone, there are ≥\geq 220 tasks, ∼40\sim 40 metrics per task, ∼10\sim 10 model families, for a total of ∼106\sim 10^{6} task-metric-model family triplets, meaning probability that no task-metric-model family triplet exhibits an emergent ability by random chance might be small. Fourthly, scientific progress can be hampered when models and their outputs are not made public for independent scientific investigation.

References

Appendix A Approximate Behavior of Metrics on Sequential Data

How do different metrics behave when used to measure autoregressive model outputs? Precisely answering this question is tricky and possibly analytically unsolvable, so we provide an approximate answer here.

Notationally, we consider NN test data of length LL (here, length is measured in tokens) with targets denoted tn=def⁡(tn1,tn2,...tnL)t_{n}\operatorname{\stackrel{{\scriptstyle\text{def}}}{{=}}}(t_{n1},t_{n2},...t_{nL}), the autoregressive model has a true-but-unknown per-token error probability of ϵ∈\epsilon\in and the model outputs prediction t^n=def⁡(t^n1,t^n2,...t^nL)\hat{t}_{n}\operatorname{\stackrel{{\scriptstyle\text{def}}}{{=}}}(\hat{t}_{n1},\hat{t}_{n2},...\hat{t}_{nL}). This assumes that the model’s per-token error probability is constant, which is empirically false, but modeling the complex dependencies of errors is beyond our scope.

Note that because we have NN test data, each of length LL, our resolution for viewing the per-token error probability ϵ\epsilon is limited by 1/NL1/NL. Here, resolution refers to “the smallest interval measurable by a scientific instrument; the resolving power.” To explain what resolution means via an example, suppose one wants to measure a coin’s probability of yielding heads. After a single coin flip, only two outcomes are possible (H, T), so the resolution-limited probability of heads is either or 11. After two coin flips, four outcomes are possible (HH, HT, TH, TT), so the resolution-limited probability of heads is now one of 0,0.5,10,0.5,1. After FF coin flips, we can only resolve the coin’s probability of yielding heads up to 1/F1/F. Consequently, we introduce a resolution-limited notation:

A.2 Token Edit Distance

We first consider an adaptation of the Levenshtein (string edit) distance for models that function on tokens rather than characters, an adaptation we term the token edit distance. The token edit distance between two token sequences tn,tn^t_{n},\hat{t_{n}} is defined as the integer number of additions, deletions or substitutions necessary to transform tnt_{n} into t^n\hat{t}_{n} (or vice versa).

The resolution-limited expected token edit distance is therefore:

From this, we see that the expected token edit distance scales approximately linearly with the resolution-limited per-token probability. The real rate is slightly higher than linear because additions and deletions contribute an additional non-negative cost, but modeling this requires a model of how likely the model is to overproduce or underproduce tokens, which is something we do not currently possess.

A.3 Accuracy

As with the Token Edit Distance (App. A.3), we ignore how likely the language model is to overproduce or underproduce tokens because we do not have a good model of this process. Continuing along,

Taking an approximation that would make most mathematicians cry:

This reveals that accuracy approximately falls geometrically with target token length. The resolution-limited expected accuracy is therefore:

From this we can see that choosing a nonlinear metric like Accuracy is affected significantly more by limited resolution because Accuracy forces one to distinguish quantities that decay rapidly.

A.4 ROUGE-L-Sum

Another BIG-Bench metric is ROUGE-L-Sum , a metric based on the longest common subsequence (LCS) between two sequences. Section 3.2 of gives the exact definition, but the key property is that ROUGE-L-Sum measures the “union” LCS, which means “stitching” together LCSs across the candidate and multiple references. As explained in the original paper: if the candidate sequence is c=w1w2w3w4w5c=w_{1}w_{2}w_{3}w_{4}w_{5}, and if there are two reference sequences r1=w1w2w6w7w8r_{1}=w_{1}w_{2}w_{6}w_{7}w_{8} and r2=w1w3w8w9w5r_{2}=w_{1}w_{3}w_{8}w_{9}w_{5}, then LCS(r1,c)=w1w2LCS(r_{1},c)=w_{1}w_{2} and LCS(r2,c)=w1w3w5LCS(r_{2},c)=w_{1}w_{3}w_{5}, then the union -LCS of c,r1,r2c,r_{1},r_{2} is w1w2w3w5w_{1}w_{2}w_{3}w_{5}, with length 4. Intuitively, this disproportionately benefits models with smaller error rates because their mistakes can be “stitched” across multiple references; this is confirmed in simulation (Fig. 9).

Appendix B Inducing Emergent Abilities in Networks on Vision Tasks

We begin by inducing an emergent classification ability in a LeNet convolutional neural network family , trained on the MNIST handwritten digits dataset . This family displays smoothly increasing test accuracy as the number of parameters increase (Fig. 10B). To emulate the accuracy metric used by emergence papers , we use subset accuracy: 1 if the network classifies KK out of KK (independent) test data correctly, 0 otherwise. Under this definition of accuracy, the model family displays an “emergent” ability to correctly classify sets of MNIST digits as KK increases from 11 to 55, especially when combined with sparse sampling of model sizes (Fig. 10C). This convolutional family’s emergent classification ability qualitatively matches published emergent abilities, e.g., at the BIG-Bench Grounded Mappings task (Fig. 10A).