A Tale of Tails: Model Collapse as a Change of Scaling Laws

Elvis Dohmatob, Yunzhen Feng, Pu Yang, Francois Charton, Julia Kempe

Introduction

Groundbreaking advances in generative AI algorithms for text, images and code are ushering in the ”synthetic data age”: increasingly we consume data generated by large scale models like GPT4 (Achiam et al., 2023), Stable Diffusion (Rombach et al., 2022) and their successors. At the same time a key driver behind the current success of large models is their consumption of massive amount of web-scale data for training. The improvements of larger models are governed by scaling laws in which error falls off as a power in the size of training data; and the emergence of new skills seems tightly linked to covering increased scales of training data. Our understanding of what the future holds in a world were models are trained on other models (or their own) synthesized data is only at its beginning, but some works indicate the possibility of complete collapse of learning, so called model collapseNot to be confused with neural collapse which refers to clustering of last-layer features at the end of training (Papyan et al., 2020).

Scaling Laws. In many domains of machine learning including speech, translation, vision, video, and mathematical problem solving, empirically observed neural scaling laws (Hestness et al., 2017; Rosenfeld et al., 2020; Kaplan et al., 2020; Hoffmann et al., 2022; Gordon et al., 2021; Henighan et al., 2021; Aghajanyan et al., 2023) demonstrate that test error often falls off as a power law with the amount of training data, model size, and compute. Theoretically, scaling laws have been derived in a variety of settings (e.g. Hutter (2021); Cabannes et al. (2023) for “LLM-like” models).

Scaling laws are intimately related to the emergence of abilities (Wei et al., 2022) in larger models, that are not present in smaller ones; and to skills that appear with decreasing loss (Gordon et al., 2021; Arora & Goyal, 2023). This bolsters the now common narrative that ”scaling is all you need”.

Current LLMs (Devlin et al., 2018; Liu et al., 2019; Brown et al., 2020; Touvron et al., 2023), including GPT-4 (Achiam et al., 2023), were trained on predominantly human-generated text; similarly, diffusion models like DALL-E (Ramesh et al., 2021), Stable Diffusion (Rombach et al., 2022), Midjourney (Midjourney, 2023) are trained on web-scale image datasets. These training corpora already potentially exhaust all the available clean data on the internet. A growing number of synthetic data generated with these increasingly popular models starts to populate the web, often indistinguishable from ”real” data. We have thus already raced into the future where our training corpora are irreversibly mixed with synthetic data and this situation stands to get worse. Recent works call attention to the potential dramatic deterioration in the resulting models, an effect referred to as ”model collapse” (Shumailov et al., 2023). Facets of this phenomenon have been demonstrated empirically in various settings (Hataya et al., 2023; Martínez et al., 2023a, b; Bohacek & Farid, 2023; Briesch et al., 2023; Guo et al., 2023). Theoretically, a few works are emerging to analyze the effect of iterative training on self-generated (or mixed) data (see Related Work): (Shumailov et al., 2023) coin the term ”model collapse” to characterize complete reversion to the mean, Alemohammad et al. (2023) analyze ”self-consuming loops” and Bertrand et al. (2023) show that iterative synthetic training leads to a ”clueless generator”.

With these first warning signs in place, we thus ask:

To this end, we carefully zoom into the scaling behavior of LLM-style models. Theoretical derivations of scaling laws always assume a heavy-tailed distribution (power-law, aka Zipf) on the input features (”heavy tail in, power scaling law out”). This distribution is of the form

Such distributions are ubiquitous in natural datasets, from Zipf’s law (Zipf, 1935) in distribution of word frequencies, to biological data, earthquake magnitudes, financial data etc. - this is the data being consumed by large models at scale, like LLMs. But what distribution do AI-models generate when trained on such data? Figure 2 provides an empirical answer for a large scale LLM (Llama2-7B) and a transformer model trained on an arithmetic task. Regenerating heavy-tailed data affects the distribution in two possible ways: (1) ”Cutting off” the tail of the distribution and/or (2) ”Narrowing” the tail (see Figure 1 for a cartoon illustration). The mechanisms leading to this, apart from finite sampling bias (as already proposed in Shumailov et al. (2023) - see Section 2 for a derivation in the Zipf-setting), stem from deliberate choices in the generation algorithms of the models: in LLMs via truncated next token prediction at inference (e.g. selecting more likely tokens via top-pp or top-kk truncation, concentrating the probability distribution by lowering the temperature); in vision models like GANs via truncation or in diffusion models through guidance.

Summary of Main Contributions. We present a high-level summary of our main theoretical contributions, some of which are highlighted in Figure 3. We empirically verify these theoretical predictions (see Figure 4): (1) in large-scale experiments on an LLM, fine-tuning Llama2-7B (Touvron et al., 2023) on an approximately 2M2M sample dataset from Wikitext-103 and (2) for transformer models trained to predict the greatest common divisor (Charton, 2023).

Assuming a true distribution as in Equation (1), consider training a model on a dataset of size TT of AI data-generated data. The synthesized data amounts to a version of the true data distribution with the tail cut at some finite rank kk or the tail narrowed to a smaller exponent. Our main findings are as follows.

(1) A Double Scaling Law.

We establish new scaling laws that explain model collapse in simplified (non-autoregressive) LM (Hutter, 2021) and toy bigram LLMs (refer to Theorems 2.1 and 4.2)The notation f(T)≲g(T)f(T)\lesssim g(T) means that f(T)≤Cg(T)f(T)\leq Cg(T) for sufficiently large TT and an absolute constant CC, while f(T)≍g(T)f(T)\asymp g(T) means f(T)≲g(T)≲f(T)f(T)\lesssim g(T)\lesssim f(T).

or equivalently (refer to Corollary 2.2), for finite-sample induced cut-off k=k(T0)k=k(T_{0}) when the generating model is trained on T0T_{0} amount of data, Etest≍T−c+T0−c′′,E_{test}\asymp T^{-c}+T_{0}^{-c^{\prime\prime}}, where the exponents c,c′,c′′c,c^{\prime},c^{\prime\prime} only depend on the tail behavior of the true distribution. This result is illustrated in Figure 3.

For AI-”tail-narrowing”, when data remains heavy-tailed, with a smaller exponent β′∈(1,β)\beta^{\prime}\in(1,\beta), the downstream Hutter LLM will scale as (Corollary 2.3)

(2) A Triplet Scaling Law for Memory-Limited Models.

We consider a simple associative memory model studied in (Cabannes et al., 2023), and establish (Theorem 5.1) a new scaling law of the form

where dd is the embedding dimension, and serves a s proxy for model capacity; the exponent cqc_{q} depends both on β\beta and the particular algorithm qq used to update the memories in the model during training.

(3) Model Collapse over Multiple Generations.

For nn-fold recursion of AI data-generation (11), where each generation of the model consumes data produced by the previous generation, we establish a universality principle of the form

where EtestcleanE_{test}^{clean} is the usual test error of the model trained on clean data (not AI-generated). This means that in Equations (2) and (4) for example, the k−c′k^{-c^{\prime}} is replaced by nk−c′nk^{-c^{\prime}}. One possible interpretation of this multiplicative degradation is that, over time (i.e as the number of generations becomes large), the effect of large language models (like ChatGPT) in the wild will be a pollution of the web to the extend that learning will be impossible. This will likely increase the value and cost of clean / non-AI-generated data.

(4) Mitigation Strategies.

In Theorem 3.2 we show that mixing AI-generated data with even a small amount of clean data mitigates model collapse by introducing a grokking phenomenon. The length of the plateau is of order kβ/πk^{\beta}/\pi, where π\pi is the proportion of training data which is from the true distribution (i.e clean data). When π=0\pi=0 (i.e only AI-generated data available), this plateau goes on forever (as in (2) and (4)). When π>0\pi>0, however small, the plateau finally halts, and the error continues to decrease à la T−cT^{-c}. This grokking phenomenon holds in the setting of deterministic ground truth labels (like in the models of Hutter (2021); Cabannes et al. (2023)). For transformer models, such deterministic settings are found for instance in arithmetic tasks, and we demonstrate it empirically in our GCD transformer experiments. The grokking effect becomes attenuated in probabilistic settings, where it can lead to an S-shaped learning curve (see Figure 19). We also identify regimes where adding AI data can be beneficial and discuss ways to curate ”tail” data to mitigate AI-data effects.

Related Work.

Theoretically, scaling laws have been derived in various settings: for non-parametric models (Schmidt-Hieber, 2017; Suzuki, 2019; Bordelon et al., 2020), in the kernel regime under the Gaussian design (Spigler et al., 2020; Cui et al., 2021, 2022, 2023; Maloney et al., 2022), or in memorization-like settings with discrete data (Hutter, 2021; Debowski, 2023; Michaud et al., 2023). Taking finite model capacity and optimization into account, Cabannes et al. (2023) recently proved scaling laws in constraint-capacity associative memories, and our Triplet Scaling Law builds on this work.

Less than a handful of works begin to provide theoretical explanations for the behavior of models in the ”synthetic data age”. (Shumailov et al., 2023) attribute model collapse to two mechanisms: a finite sampling bias cutting off low-probability ”tails”, thus leading to more and more peaked distributions and function approximation errors; they theoretically analyze the (single) Gaussian case and provide empirical evidence for VAEs, Gaussian mixtures and the OPT language model (125M parameters). In the context of vision models, Alemohammad et al. (2023) analyze ”self-consuming loops” by introducing a sampling bias that narrows the variance of the data at each generation, and, in addition to empirical demonstration on GANs and denoising diffusion probabilistic models, provide theoretical analysis for the Gaussian model. Finally, let us mention the study of Bertrand et al. (2023) which sheds light on the critical role of data composition in the stability and effectiveness in generative models, applicable to VAEs (Kingma & Welling, 2014), diffusion models and normalizing flows. They explore scenarios involving a mix of clean data, representative of the true distribution, and synthesized data from previous iterations of the generator. Their analysis reveals that if the data mix consists exclusively of synthesized data, the generative process is likely to degenerate over time (”clueless generator”). Using fixed-point analysis across iterations, they find that when the proportion of clean data in the mix is sufficiently high, the generator, under certain technical conditions, retains the capability to learn. A recent paper (Fan et al., 2023) empirically observe deteriorated scaling laws when training on synthetic data for text-to-image models. A more detailed description of related and prior work can be found in Appendix A

To our knowledge, our work is the first to theoretically and empirically analyze model collapse in the context of scaling laws and emergent abilities to provide a rich new landscape of AI-data induced phenomena.

A Deterministic Infinite Memory Model

Here, we present the core of our theory for the simplest case of (i) infinite memory and (ii) a deterministic ground truth labeling function i↦yii\mapsto y_{i}, studied by Hutter (2021) (the ”Hutter LLM”). Both restrictions will be lifted in later sections, where we also analyze an probabilistic autoregressive version (Section 4) and limited memory models (Section 5). Token ii is drawn according to the Zipf law in Equation (1), which e.g. models distribution of various metrics in language. Another interpretation of the appearance of a power-law is offered by the ”Quantization Hypothesis” paper of Michaud et al. (2023): one may think of each ii as some discrete skill, needed to solve a problem for example; thus, the skills occur at different rates pip_{i}. The shape parameter β>1\beta>1 controls the length of the tail of this distribution: bigger values of β\beta correspond to longer tails.

As mentioned, deliberate choices in the AI generation algorithm (like top-pp or top-kk next token prediction) immediately lead to a chopped tail at kk. When viewed as skills, we can say that only the kkth most frequent outcomes (”skills”) are considered. But even when no tails are cut deliberately, the finite size T0T_{0} of the training set (sampling bias) induces an effective tail-cutting. This can be seen as follows: Sample an iid dataset of size T0T_{0}, and estimate the histogram pAIp_{\text{AI}}; this new distribution plays the role of an AI data-generator. An integer ii appears in the support of pAIp_{\text{AI}} a number of times which is T0piT_{0}p_{i} on average. Roughly speakingThis can be made rigorous via standard concentration arguments., this means that the support of pAIp_{\text{AI}} is {i∣pi≤C/T0}={i∣i≤k}\{i\mid p_{i}\leq C/T_{0}\}=\{i\mid i\leq k\}, where

Therefore, the transformation p→pAIp\to p_{\text{AI}} amounts to chopping off the tail of pp at rank kk, where kk is as given above.

Tail Narrowing. Figure 2 (for Llama2) shows that in addition to tail cutting, tail narrowing effects happen during AI-generation. One mechanism for this is lowered temperature during next-token prediction. Assume a softmax distribution on the logits ziz_{i} for the iith token: pi=ezi/∑jezjp_{i}=e^{z_{i}}/\sum_{j}e^{z_{j}}. Define qiT=ezi/T/∑jezj/Tq_{i}^{T}=e^{z_{i}/T}/\sum_{j}e^{z_{j}/T} for general temperature TT. Then pi≍i−βp_{i}\asymp i^{-\beta} morphs into qiT≍i−β/Tq_{i}^{T}\asymp i^{-\beta/T} (to first order). We see that temperature scaling directly causes narrowing of tail for T>1T>1. Other mechanisms can come to play: for instance, for autoregressive models with perplexity, token-wise tail cutting can result in tail narrowing for sequence-perplexity (see Figure 35 and discussion in Appendix I).

2 A New Scaling Law in the Hutter LLM

For a deterministic ground-truth labelling function i↦jii\mapsto j_{i}, consider a downstream Hutter “LLM” (Hutter, 2021)

constructed on a sample DT:={(it,jt)∣t∈[T]}\mathcal{D}_{T}:=\{(i_{t},j_{t})\mid t\in[T]\} of size TT from unmitigated Zipf distribution pp (1), the test error obeys the following scaling law Hutter (2021)

Now, let qq be a k-tail-cutting version of pp, i.e qi∝piq_{i}\propto p_{i} if i≤ki\leq k and qi=0q_{i}=0 otherwise. When constructed (“trained”) on DT\mathcal{D}_{T} of size TT, now from qq, the test error (w.r.t to the true data distribution pp) of this model is

That is, we train on data from the AI distribution qq and test on original distribution pp. We prove the following scaling law for tail cutting (all proofs are relegated to Appendix C):

Consider long-tail real-world data with exponent β>1\beta>1, and let the cutoff for AI-generated data be kk. Then, for large kk and TT samples from the AI, the test error of the downstream ”LLM” scales like so E_{test}\asymp T^{-(\beta-1)/\beta}+{\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}k^{-(\beta-1)}}\asymp\min(T,{\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}k^{\beta}})^{-(\beta-1)/\beta}.

Thus, as soon as T≳kβT\gtrsim k^{\beta}, the AI-generated sample size TT ceases to be a ”scalable” resource: collecting more AI-generated samples will not improve the performance of the downstream model, i.e performance plateaus and we lose scaling. The result is illustrated empirically in Figure 3, left and Figure 9 (Appendix B).

When we assume that the AI-generator itself was trained on T0T_{0} samples, we get a similar loss of scaling stemming from the tail cutting from finite sampling bias (Equation (6)):

These theoretical are empirically confirmed in the Figure 9.

In the case of tail narrowing, the scaling behavior changes; instead of a plateau, we obtain a slower decay rate:

In the setting of Theorem 2.1, consider AI-generated data to also be long-tail data, albeit with smaller exponent β′∈(1,β)\beta^{\prime}\in(1,\beta). Then, the downstream Hutter LLM trained on AI-generated data will scale as Etest≍T−(β−1)/β′E_{test}\asymp T^{-(\beta-1)/\beta^{\prime}}.

3 Collapse Over Multiple Generations of AI Data

We now examine the cumulative impact of prior loss of scaling across multiple generations. Consider nn-fold recursive AI data-generation, i.e

Each arrow corresponds to drawing a sample of size T0T_{0}. If we iterate nn times the argument leading to (10), we get the following scaling for the test error Etest(n)=Etest(n)(T)E_{test}^{(n)}=E_{test}^{(n)}(T) for learning on TT samples from the nnth generation and testing on the true data distribution,

where c:=1−1/βc:=1-1/\beta. We deduce the following result.

Model collapse (as spoken of in the literature) occurs iff n≫(T0/T)cn\gg(T_{0}/T)^{c}.

For example, if T0≫TT_{0}\gg T (e.g T0≥CTlog⁡TT_{0}\geq CT\log T) and nn is constant (e.g n=25n=25), then model collapse will not occur if we learn on the nnth generation of AI data. On the other hand, if T0≲TT_{0}\lesssim T, then model collapse will eventually occur.

In particular, taking T0≍TT_{0}\asymp T, we get

Note how the loss scales linearly with the number of generations. Figure 3, middle, illustrates how an increased number of generations moves the loss scaling curve progressively to the right. This leads to eventual model collapse.

Mitigating Model Collapse via Data Mixing

Here we explore the possibility of alleviating model collapse via the acquisition of even a tiny amount of data from the true data distribution, to complement AI polluted data. We study two phenomena: (1) In the case of mixing π\pi-fraction of the original data with a (1−π)(1-\pi) fraction of AI-generated data we exhibit a startling ”grokking” phenomenon where test loss plateaus with increasing training data to finally decrease again according to the scaling law of the original model, and (2) in the scenario where we would like to compensate for missing ”tail”, we acquire some data from the tail of the original distribution to show that this needs to be done with caution: getting data from ”too deep” in the tail is worthless while data closer to the precise ”missing” tail can be beneficial. All proofs can be found in Appendix C.

To counter the effect of tail cutting and the resulting plateau in scaling, we might resort to adding curated data that would emphasize the tail. The following Theorem 3.1 studies this effect; it shows, in particular that if we ”overshoot” and only curate tail that is too deep, our efforts will be worthless. Rather, there is a fine line around the chopped tail kk (within a factor of (1+o(1))(1+o(1)) of kk), where we need to place our data curation efforts to achieve the desired effect, a return of the scaling beyond the plateau.

Suppose we “buy” a chunk of the tail of the real data distribution corresponding to i=N,N+1,…i=N,N+1,\ldots; let the distribution be π\pi (thus, supported on {N,N+1,…}\{N,N+1,\ldots\}). Now, let kk, NN, and TT tend to infinity such that N/k→CN/k\to C, with C∈[1,∞]C\in[1,\infty]. We have the following sharp phase-transition.

(A) If C=1C=1, e.g if N=k+kN=k+\sqrt{k}, then Etest≍T−cE_{test}\asymp T^{-c}. That is, we perfectly anneal the tail-chopping effect of AI-generated data.

(B) If C>1C>1, then Etest≍T−c+k−αE_{test}\asymp T^{-c}+k^{-\alpha} (which recovers the result of Theorem 2.1), and so ”buying” the NNth tail of the real data distribution is worthless.

2 A Grokking Phenomenon

Here we show how even small amounts of original data can mitigate the above ”scaling law collapse” by introducing a grokking phenomenon where test error plateaus and eventually continues to decline.

Consider a sample of size TT of which a proportion π\pi comes from the true distribution pp and the remainder comes from a version p′p^{\prime} of pp with its tail chopped off at rank kk. We have the following scaling laws for the Hutter LLM define din (7).

(A) Early-Stage Dynamics. For T≪kβ/πT\ll k^{\beta}/\pi, it holds that

Thus, during this stage, the money spent on acquiring some clean data is not amortized!

(B) Later-Stage Dynamics. As soon as T≥Ckβ/πT\geq Ck^{\beta}/\pi (where CC is an absolute constant), it holds that

Thus, during this stage, we recover the unpolluted sample-size law scaling T−(1−1/β)T^{-(1-1/\beta)}, up to within a multiplicative constant π−(1−1/β)\pi^{-(1-1/\beta)} (which can be seen as an increase in the price of data). For fixed TT and tunable π\pi, this error rate scales like π−(1−1/β)\pi^{-(1-1/\beta)}, which is yet another scaling law.

Effectively, the above theorem predicts that for any fixed π∈(0,1)\pi\in(0,1) –no matter how small– the test error grokks w.r.t sample size TT. The result is empirically confirmed in Figure 3, right (see Figure 10 for another illustration).

We experimentally confirm this new phenomenon for transformer models trained to calculate the GCD (see Appendix G), which indicates its applicability for a wider class of LLMs with underlying deterministic ground truth, like for arithmetic tasks.

In Appendix C.4 we state and prove a similar theorem in the case of tail narrowing of synthetic data.

The above machinery allows us to analyze a particular regime where AI-data can help improve performance.

Taking T=Treal+TAIT=T_{real}+T_{AI} and π=Treal/T\pi=T_{real}/T, we have the following important corollary of Theorem 3.2.

For Treal≪kβT_{real}\ll k^{\beta}, it holds that Etest≍(Treal+TAI)−(1−1/β)+k−(β−1).E_{test}\asymp(T_{real}+T_{AI})^{-(1-1/\beta)}+k^{-(\beta-1)}.

Figure 5 illustrates how AI data can boost performance, up to a certain point, when its benefits plateau. This result might contribute to our understanding of why, sometimes, adding AI-generated data might lead to better models, especially when generated by a stronger model (e.g. He et al. (2023); Shipard et al. (2023); Bansal & Grover (2023); Lin et al. (2023)). See Appendix A for more references.

A Tailed Bigram Model

We will now proceed to a more complex model, bringing us closer to capturing the probabilistic and autoregressive nature of LLMs (next token prediction). In this Section we will define the data generating process, define the new model (Hutter++), and establish that the original scaling law (with clean data) still holds. We then proceed to show similar loss of scaling for AI-data.

(instead of j−βj^{-\beta}), where πi\pi_{i} is a permutation associated to every ii providing the order of outputs. To summarize, we think of the data as pairs (i,j)(i,j), where the distribution of ii is governed by some p(i)p(i) as in the deterministic Hutter setting, and p(j∣i)p(j|i) is given by Equation (16).

This setting can be made autoregressive by generating sequences step by step, using the preceding output as the next input. We can think of each successive pair of tokens as of the pairs (i,j)(i,j) above, with the only difference that the marginal distribution p(i)p(i) changes. We thus will make no assumptions on p(i)p(i) in what follows (except for a mild technical condition). Proofs can be found in Appendix D.

We now present an extension of the Hutter model (7) which is adapted to bigrams. Let nT(i)=∑t=1T1[it=i]n_{T}(i)=\sum_{t=1}^{T}1[i_{t}=i] be the number times the context iti_{t} appears in the dataset DT\mathcal{D}_{T} and nT(i,j)=∑t=1T1[(it,jt)=(i,j)]n_{T}(i,j)=\sum_{t=1}^{T}1[(i_{t},j_{t})=(i,j)] be the number of times the pair (i,j)(i,j) appears in the dataset. Note that nT(i)∼Bin(T,pi)n_{T}(i)\sim Bin(T,p_{i}). As soon as nT(i)≥1n_{T}(i)\geq 1, define

This is an empirical version of p(⋅∣i)p(\cdot\mid i) based on an iid sample of size nT(i)n_{T}(i). For a theoretical analysis, we shall consider the following test error metric based on total-variation (TV)

2 A Scaling Law for Hutter++

Suppose β∈(1,∞)∖{2}\beta\in(1,\infty)\setminus\{2\} and set c:=min⁡(1−1/β,1/2)c:=\min(1-1/\beta,1/2). If ∑ipi1−c<∞\sum_{i}p_{i}^{1-c}<\infty, then Etest≲T−c.E_{test}\lesssim T^{-c}. Moreover, if β∈(1,2)\beta\in(1,2) and the mappings π1,π2,…\pi_{1},\pi_{2},\ldots are permutations, then Etest≍T−cE_{test}\asymp T^{-c}.

Thus, the proposed Hutter++ algorithm induces exactly the same scaling law as the classical setup (Hutter, 2021) !

3 Model Collapse in Probabilistic Setting

We now return to our main problem, understanding model collapse in the probabilistic setting and consider the Hutter++ presented above. Thus, suppose the learner only has access to at most a dataset of size TT containing the kkth head of the conditional distribution p(⋅∣i)p(\cdot\mid i). That is, sampled from: i∼pi\sim p, j∼p(j∣i)1[j≤k]j\sim p(j\mid i)1[j\leq k] (normalized appropriately), where p(⋅∣i)p(\cdot\mid i) is as in Equation (16).

(A) If β∈(1,∞)∖{2}\beta\in(1,\infty)\setminus\{2\} and ∑ipi1−c<∞\sum_{i}p_{i}^{1-c}<\infty where c:=min⁡(1−1/β,1/2)c:=\min(1-1/\beta,1/2) as before, then Etest≲T−c+k−βc.E_{test}\lesssim T^{-c}+k^{-\beta c}.

(B) Furthermore, if the mappings π1,π2,…\pi_{1},\pi_{2},\ldots are permutations and ∑ipi1−c<∞\sum_{i}p_{i}^{1-c}<\infty, then Etest≍T−c+k−βcE_{test}\asymp T^{-c}+k^{-\beta c}.

Multiple Generations. The mechanics of the proof of Theorem 2.4 apply in this setting. See Figure 12 in Appendix B illustrating that Equation (13) keeps holding for probabilistic data distributions.

Grokking for Mixtures. Technically speaking, this grokking phenomenon only holds for models with deterministic ground truth labels, like the Hutter LLM and the limited capacity associative memory model. For the probabilistic setting of bigrams (or text LLMs) the theorem cannot hold in its pure form, because if we train on a mixture of two distributions (clean and synthetic) but test only on the clean distribution, the distance between these two distributions will always be a lower bound on the test error. However, we can see that remnants of a ”smoothed” grokking-law persist in the form of an S-shaped scaling (see Figure 19 in Appendix B).

Capacity-Limited Memory Models: A Triplet Scaling Law

We now turn to a finite-memory extension of the Hutter LLM, which allows to model capacity. We thus study the model collapse phenomenon in the context of the following simple associative memory model studied in (Cabannes et al., 2023)

This is a transformer-like finite-memory extension of the infinite-memory model in (Hutter, 2021). The integer d≥1d\geq 1 then plays the role of the ”capacity” of the resulting model. Here, f⋆:[N]→[m]f_{\star}:[N]\to[m] is an unknown function, for example, reduction modulo mm, i.e f⋆(i):=((i−1) mod m)+1f_{\star}(i):=((i-1)\text{ mod }m)+1; qT=q(DT)q_{T}=q(\mathcal{D}_{T}) is probability distribution on [N][N] which encodes an arbitrary learner, estimated using and iid sample Dt={(it,yt)∣t∈[T]}\mathcal{D}_{t}=\{(i_{t},y_{t})\mid t\in[T]\} of size TT collected from a probability distribution on [N]×[m][N]\times[m], of the form

In the context of model collapse which is the main focus of this manuscript, we have the following.

For all the algorithms qq considered in (Cabannes et al., 2023), one has the following triplet scaling law w.r.t sample size TT, embedding dimension dd, and frequency cutoff kk,

This result is empirically confirmed in Figure 21 and proved in Appendix E. It gives rise to similarly tapered-off scaling curves for synthetic data, as in the simpler models. The proofs for loss of scaling across generations in Section 2 and grokking phenomena in Section 3, carry over to this model as well, demonstrating their universality.

Experiments

In this section we present our experimental results to demonstrate evidence of various predictions we have made theoretically. We showcase four scenarios of increasing level of complexity: an empirical Hutter++ model, autoregressive bigram models with perplexity loss, an arithmetic transformer to predict the GCD of two integers (Charton, 2023) and a large-scale LLM, Llama2-7B (Touvron et al., 2023), trained on a large data corpus (Wikidata-103).

In our theoretical analysis, motivated by empirical observations (see Figure 2) or by the effect of finite data sampling bias on heavy-tailed data, we have assumed that generated data follows patterns of either a tail cutoff or tail narrowing. In our subsequent experiments, we depart from theoretical assumptions on tail-cutting/narrowing to allow the widely deployed top-p selection or temperature scaling mechanisms to give possibly intermingled effects on the generated data distribution.

Empirical Hutter++ Model. In Figure 6, we use an initial model that is trained on T0=100,000T_{0}=100,000 samples from the original distribution. For the Gen 1 line, the data are all generated from this initial model. From Gen 2 onwards, models are iteratively trained on data produced by the most performant model of the preceding generation, effectively eliminating the possibility that model collapse results from inadequate sampling. For Gen 1, a notable degradation in data scalability is observed, alongside a rapid decline in model performance across generations. These observations not only validate our theoretical result but also reaffirm our assumptions. A similar pattern is evident with temperature scaling, as shown in Figure 16.

Autoregressive Bigram Models with Perplexity Loss. We move one step further towards ”real” LLMs to investigate autoregressive bigram models. The dataset now comprises sequentially generated integers, adhering to Equation (16), with the model trained on all tokens. We use the averaged perplexity score of the test set as the test error metric. Our study encompasses a range of effects—such as top-p inference, temperature scaling, limited real data, and training on progressively larger AI datasets. Consistent with the findings in Section 4, we observe the same patterns of scaling loss and progressive model collapse across generations. Relevant figures are provided in Appendix F.

Transformers Learning the GCD. Our first illustration of our theory ”in the wild” is for sequence-to-sequence transformer models for an arithmetic task: predicting the greatest common divisor (GCD) of two integers, encoded as sequences of digits in some base BB following Charton (2023). This setup is a perfect intermediate step between our toy models and large scale LLMs; it uses the transformer architecture and training algorithms on sizeable models, while the underlying data has a deterministic nature. Over the course of training the model progressively learns new GCDs and with them also their products with already learned GCDs. We can thus view each such learned group, usually learned in ”bursts”, as a new skill. For the purpose of this experiment, we use this model after 300M300M samples as the generator of AI-data. In Figure 4 we validate the predicted scaling law for a single generation and observe ‘un-learning’ of skills when training exclusively with generated data, as well as a grokking effect when training with mixtures. See Appendix G for a full description and more details and Figures.

Experiments on LLMs. We finetune Llama2 with LoRA, generating synthetic AI data for the next finetuning iteration. Inspired by the setup in Shumailov et al. (2023), we use Wikidata-103, partitioned into approximately 2.22.2 million sequences of 128 tokens. AI data is generated through prompt completion, using the first 96 tokens from the original sequences as prompts. The model is trained only on the last 32 tokens to preclude information leakage, i.e. the model being trained on the ground truth of the same 32 tokens. The evaluations are conducted exclusively on the same 32 tokens. We use top-p 0.9 and temperature 0.9 across all generation. The results, depicted in Figure 4 (left), illustrate a scaling law decay over several generations. The first generated dataset still contain useful but limited information and the utility of the second generation’s data markedly diminishes. These phenomena corroborate the anticipated loss of scaling law and model collapse, further indicating that model collapse is even more pronounced here, highlighting the challenges in training next generation LLMs. More details and results in Appendix H.

Moreover, we conduct experiments to investigate mixing a proportion of real data with AI-generated data. Figure 7 demonstrates the effect of blending a random 2% of original data with AI data across all fine-tuning phases. It significantly mitigates model collapse, with the emergence of a grokking curve as predicted in Theorem 3.2.

Conclusion

In the advent of the ”synthetic data age”, our work signals the end of current neural scaling laws and opens the door to a puzzling array of new phenomena in a world where the training corpora are enriched with AI generated data. We demonstrate that scaling laws cease to persist; test error tapers off due to altered, less heavy tailed, data distributions. As noted already in prior work, in a fully synthetic data world learning will stop and models will degenerate - their scaling will halt and revert completely. Yet, new opportunities arise from careful mixture and data curation, as we have shown, with interesting effects at the interplay of clean and synthesized data. We must recognize new learning plateaus and, for instance, adjust to changed learning curves from blending clean and synthetic data to unintended early stopping. A notable feature of our work is that our theory is effective - we observe the predicted phenomena for relevant large models in two different settings.

Taken together, our contributions call for a more responsible, or ”collapse-aware”, proliferation of synthesized data. Scale is not all you need: more work on effective watermarking for synthetic data is needed, to make it more distinguishable from the original, human-annotated data. Thus, clean / real data will become an even more valuable resource in the future, as we are ushering in the ”beyond scaling” era.

Acknowledgements

YF and JK acknowledge support through NSF NRT training grant award 1922658. YF and PY would like to thank Di He for discussions and suggestions. This work was supported in part through the NYU IT High Performance Computing resources, services, and staff expertise.

References

Appendix A Prior Work

The phenomenon appeared in the recent literature in the context of language and image generation. Several recent works demonstrate facets of this phenomenon empirically in various settings (Hataya et al., 2023; Martínez et al., 2023a, b; Bohacek & Farid, 2023; Briesch et al., 2023; Guo et al., 2023; Fan et al., 2023). Only few recent works also provide some accompanying theoretical analysis (Shumailov et al., 2023; Alemohammad et al., 2023; Bertrand et al., 2023) which we outline now.

(Shumailov et al., 2023) define model collapse and attribute it to two mechanisms: finite sampling when training a model (leading to cut off of low-probability data) and function approximation errors (the model is not sufficiently expressive to model the true distribution). They observe (and, for a single Gaussian, prove) that upon iteratively resampling finite ”training data” the generated distribution becomes more and more peaked. Other models studied empirically are mixtures of (two) Gaussians and VAEs on MNIST. To study language models, (Shumailov et al., 2023) iteratively fine tune Meta’s OPT-125M model on wikidata2. For generation of new text they use a 5-way beam search, which, by its nature, (approximatively) generates only low-perplexity data.

(Alemohammad et al., 2023) conduct an empirical and analytical analysis on generative image models of what they term the ”self-consuming” or ”autophaguous” loop. They conclude that without enough fresh real data at each generation, future models necessarily will have their precision or recall decrease. They model the influence of each new AI-generation via a generic sampling bias 0≤λ≤10\leq\lambda\leq 1. In the case of image generation this refers to feature parameters at generation that favor quality over diversity (suitably quantified). More precisely, λ=1\lambda=1 corresponds to unbiased sampling and λ=0\lambda=0 corresponds to sampling from the modes of the generative distribution. λ\lambda models biased sampling methods commonly used in generative modeling practice, such as truncation in BigGAN and StyleGAN or guidance in diffusion models. In the case of Gaussian distributions, λ\lambda is the shrinking factor of the variance of the next generation. Their empirical work studies GANs and denoising diffusion probabilistic models for image generation on FFHQ and MNIST and single Gaussians for both theoretical and empirical observations. As in (Shumailov et al., 2023) they observe (and prove for the case of a single Gaussian) that estimation error alone leads to vanishing variance with number of iterations. (Alemohammad et al., 2023) also empirically observe an initial boost in performance in a regime where modest amounts of synthetic data are mixed with the original data before larger amounts of synthetic data lead to ultimate degradation. This might mimick larger-scale results that demonstrate how synthetic data mixed with true data improves performance in some scenarios (see Benefits of synthesized data below). Indeed, in its simplest form, data augmentation (rotations, cropping etc. ), a widespread highly beneficial practice in ML training, can be viewed as the simplest form of data generation.

Let us mention the study of Bertrand et al. (2023) in the context of image generation, which sheds light on the critical role of data composition in the stability and effectiveness in generative models. They explore scenarios involving a mix of clean data, representative of the true distribution, and synthesized data from previous iterations of the generator. Their analysis reveals that if the data mix consists exclusively of synthesized data, the generative process is likely to degenerate over time, leading to what they describe as a ’clueless generator’. Thus, the generator collapses: it progressively loses its ability to capture the essence of the data distribution it was intended to model. Conversely, they found that when the proportion of clean data in the mix is sufficiently high, the generator, under certain technical conditions, retains the capability to learn and accurately reflect the true data distribution. This work sheds light on the critical role of data composition in the stability and effectiveness of generative models.

Several empirical studies confirm the deleterious effect of training on self-generated data: In the context of image generation, (Martínez et al., 2023a, b) report degradation of models trained on AI-generated data. Specifically, they use a Denoising Diffusion Implicit Model and a few (relatively small) datasets (e.g. Orchids, MNIST) to demonstrate visual degradation when training in successive generations of AI-generated data. (Hataya et al., 2023) ”conclude that generated images negatively affect downstream performance, while the significance depends on tasks and the amount of generated images”, (Bohacek & Farid, 2023) reports that the popular StableDiffusion model collapses when iteratively retrained on self-generated faces, even with as little as 3%3\% synthetic data mixed into the original training set. For text, (Briesch et al., 2023) use nanoGPThttps://github.com/karpathy/nanoGPT on a curated 10K logical-expression dataset to demonstrate the iterative collapse of self-consuming loops - the model and dataset are sufficiently small to allow training from scratch. (Guo et al., 2023) observe a decline in linguistic diversity metrics across iteratively fine-tuned LLMs.

Mitigation:

To our knowledge, rigorous theory (or even empirical demonstrations) on mitigation strategies against model collapse are yet to come, with one notable exception in (Bertrand et al., 2023) (see below). Several works discuss the need for detection of AI-generated images or text (to avoid retraining on them), for example motivating research into watermarking strategies. (Bertrand et al., 2023) analyze iterative retraining on a mixture of synthesized and original data under several technical assumptions and find that there are fixed points governing the stability of iterative retraining.

Benefits of Synthesized Data

There is a range of results showing benefits of AI-synthesized data in training better models, though mostly these results pertain to image data, specifically in the context of diffusion models (Azizi et al., 2023; He et al., 2023; Shipard et al., 2023; Bansal & Grover, 2023; Lin et al., 2023), though not only (see Dai et al. (2023); Xu et al. (2023); Huang et al. (2022); Wang et al. (2023) for chat-related examples). One might argue that they either throw model-collapse caution to the winds or, possibly, settle in the protected corner where mild amounts of synthetic data (or larger amounts of ”mildly synthetic” data, like in the case of data augmentation) helps. In particular, often benefits of synthetic data are observed when the synthetic data is generated by a model trained for a different use case than the downstream task (like images synthesized from diffusion models helping classification models) or generated by a stronger model (He et al., 2023; Shipard et al., 2023; Bansal & Grover, 2023; Lin et al., 2023). However, other works critically analyze the purported benefit of generated data. (Burg et al., 2023) find that while synthesized data from a diffusion model helps improving downstream tasks, such as classification, using the pre-training data of the diffusion model alone gives even stronger performance (which we can interpret as evidence of mild first-generation model collapse). All in all it is fair to say that the impact of data augmentation using generative models is still not fully understood.

Scaling Laws:

Neural scaling laws have been ubiquitously observed in vision, language and speech. Early large scale empirical studies are performed in (Hestness et al., 2017; Rosenfeld et al., 2020), demonstrating power law scaling across a range of learning scenarios. This is followed by well-known large-scale studies from OpenAI (Kaplan et al., 2020) and DeepMind (Hoffmann et al., 2022), which empirically demonstrate power-law scaling in LLMs across a wide set of scales. Essentially, this empirically establishes that

where LL is the per-token cross entropy loss (in nats), N,DN,D are the number of (non-embedding) parameters and data, respectively, and NC,DCN_{C},D_{C} and αN,αD\alpha_{N},\alpha_{D} are constants determined by the data distribution and the model specifications.

This study was extended to demonstrate many more power law relations in various scenarios (vision transformer, video modeling, multimodal models, and mathematical problem solving) (Henighan et al., 2021). In the machine translation (MT) setting, (Gordon et al., 2021) quantify scaling laws for standard benchmarks like BLEU and explain them via cross-entropy power-law scaling, thus positing a first universality of scaling laws across metrics. (Hernandez et al., 2021) demonstrate similar empirical power-law scaling for transfer learning and (Aghajanyan et al., 2023) provide a vast experimental body of evidence for scaling laws in mixed-modal language models.

However, a few results have nuanced the view of scaling as a panacea to improved loss. For instance, (McKenzie et al., 2023) present evidence for ”inverse sclaing” where flaws in the training objective or the data lead to U-shaped scaling.

Theoretical Models for Scaling Laws:

From a theoretical angle, scaling laws have been shown analytically even before the emergence of large foundation models. For instance, Caponnetto & de Vito (2007) characterize the power-law generalization error of regularized least-squares kernel algorithms. The role of optimization can also be taken into account in this setting ((Nitanda & Suzuki, 2021)). In the nonparametric literature, for example (Schmidt-Hieber, 2017) and (Suzuki, 2019) derived the test error scaling of deep neural network in fitting certain target functions and (Bordelon et al., 2020) analyze spectral dependence.

More recently, scaling laws have been shown for kernel models under the Gaussian design, e.g. in (Spigler et al., 2020; Cui et al., 2021, 2022) for regression and (Cui et al., 2023) for classification. (Maloney et al., 2022) study scaling laws for the random feature model in the context of regression. In the context of memorization for heavy-tailed data scaling laws have been shown in the infinite-memory setting (Hutter, 2021), for ”quantized” skills (Michaud et al., 2023) and for certain random data-generation processes (Debowski, 2023). When taking model capacity and optimization into account, Cabannes et al. (2023) recently proved scaling laws in constraint-capacity associative memories.

To our knowledge, however, very few papers deal with the decay of scaling in the case of self-consuming loops. A notable example is (Mobahi et al., 2020) which studies iterated retraining in the context of self-(knowledge-)distillation in the kernel setting. However, this analysis is very distinct from our work, not only because it places itself in the kernel setting with Gaussian design, but also because it assumes the distillation setting, where the ”generation” stage is carefully optimized for the next stage training. In the case of synthesized data in the wild, this assumption can of course not be made.

Emergence of “Skills” and Scaling Laws:

Scaling laws give us an insight on bang-for-the-buck style trade-off for model training. However, cross-entropy loss is not a goal in and of itself: we want to train models that are endowed with a larger and larger skill set as we scale them up. For instance, (Gordon et al., 2021) provide intuition and empirics for the scaling of BLEU score for MT with cross-entropy loss as

demonstrating “emergence” of good BLEU performance with scale. This type of “emergence” has been massively confirmed in (Wei et al., 2022), where a working definition of “emerging” is “not present in smaller models, but appears in larger models”. In this sense, (Wei et al., 2022) demonstrate empirically a large number of “skills” appearing with scale, like Multi-Task NLU, Modular arithmetic, word unscrambling and transliteration.

A theoretical model, providing an underpinning of the necessity of scaling laws for the emergence of skill has recently been given by (Arora & Goyal, 2023). They analyse “emergence” with the scaling laws as a departure point in a model that links cross-entropy loss in LLMs to basic skills to show that scaling laws enable the model to learn (and generalize) efficiently.

Strengthening the tie between scaling laws and emergent skill, albeit in the opposite direction, (Michaud et al., 2023) posit that skills that emerge in ”quanta” imply a scaling law of the loss. Related, (Chen et al., 2023) assume a hierarchy of skills to derive data curation mechanisms to precipitate the emergence of skills, though they do not allude to scaling laws directly.

Appendix B Complimentary Figures for Sections 2, 3 and 4

Figures 9, 9 and 10 further illustrate our theory for simple Hutter LLM.

Hutter++.

We now provide complementary illustrations of predictions made from the theory we have developed for the generalized Hutter models as in Equation (16) in Section 4, without departing from our theoretical assumptions. We also show how theory from the infinite memory model in Section 2 continues to hold in this bigram setting. Figure 12 confirms the scaling law of Theorem 4.2.

In Figure 3 (middle) we have seen an illustration of the translated scaling curves under n-fold synthesized data in the Hutter LLM. Figure 12 illustrates this phenomenon for the slightly more complex tailed bigram model.

Both Figures 3 (middle) and 12 illustrate the setting where each model consumes as much training data as its predecessor (T0=TT_{0}=T). We now relax the assumption that each successive model has strictly the same amount of training data as its predecessor. We assume that the generation 0 model is trained on T0T_{0} (here, T0=100,000T_{0}=100,000) amount of original data to generate AI data for generation 1. All future generations, starting from generation 2, are trained on data generated by the most powerful model from the previous generation (T=1,000,000T=1,000,000 data in this case). Figure 14 (for Hutter LLM) and 14 (for Hutter++ on paired bigram data) show the resulting scaling behavior. We take this setting even further by adding a top-p tail cutting mechanism and a temperature scaling mechanism for each synthetic data generation. Figure 6 cuts at p=0.95p=0.95 and Figure 16 at temperature 0.90.9.

We now study mixing of clean and synthesized data in the bigram setting. Figures 18 and 18 add top-p tail-cutting when synthesizing, and start with T0=10,000T_{0}=10,000 original data samples, which are successively blended with synthesized data from the largest model. Note that in this setting we observe a reversion of scaling laws with increased AI data. This needs to be compared with the orange curve in Figure 20 in the deterministic Hutter setting. The probabilistic nature of the bigram models leads to a new effect here.

Appendix C Proofs for the infinite memory (Hutter) Model (Sections 2 and 3)

Observe that the model f^\widehat{f} makes an error on ii if and only if the iith ”skill” never occurred in the training dataset DT\mathcal{D}_{T}, i.e either (1) i≥k+1i\geq k+1, or (2) 1≤i≤k1\leq i\leq k and it≠ii_{t}\neq i for all t∈[T]t\in[T]. We deduce that

where c:=1−1/β∈(0,1)c:=1-1/\beta\in(0,1), and we have used the elementary fact that ∑i≥k+1i−β≍k−(β−1)\sum_{i\geq k+1}i^{-\beta}\asymp k^{-(\beta-1)} for large kk. For the second sum, we will need the following lemma.

Consider the function h(z):=ze−Tzh(z):=ze^{-Tz} for z∈(0,1)z\in(0,1). Its derivative is h′(z)=e−Tz(1−Tz)h^{\prime}(z)=e^{-Tz}(1-Tz). Thus, hh is increasing on (0,1/T)(0,1/T) and decreasing on (1/T,∞)(1/T,\infty). Furthermore, note that pi≤1/Tp_{i}\leq 1/T iff i≥T1/βi\geq T^{1/\beta}. We deduce that

For the second part, note that Γ(c,T)=o(1)\Gamma(c,T)=o(1) for large TT so that

We now consider two separate cases for the relative scaling of kk and TT.

since k−(β−1)≳T−(β−1)/β=T−ck^{-(\beta-1)}\gtrsim T^{-(\beta-1)/\beta}=T^{-c}.

Here, thanks to Lemma C.1 we have Γ(c,T)=o(1)\Gamma(c,T)=o(1) and Γ(c,Tk−β)=Θ(1)\Gamma(c,Tk^{-\beta})=\Theta(1). We deduce that

since k−(β−1)≲T−(β−1)/β=T−ck^{-(\beta-1)}\lesssim T^{-(\beta-1)/\beta}=T^{-c}. Putting things together then gives the claimed result. ∎

C.2 Proof of Corollary 2.3

Indeed, let pi∝i−βp_{i}\propto i^{-\beta} and (pAI)i=qi∝i−β′(p_{AI})_{i}=q_{i}\propto i^{-\beta^{\prime}}. Then,

That is, Etest≍T−cE_{test}\asymp T^{-c} as claimed. ∎

C.3 Proof of Theorem 3.2 and Corollary 3.3

Suppose that of TT samples available for training our model, πT\pi T are samples from the true distribution p=Zipf(β)p=Zipf(\beta) and (1−π)T(1-\pi)T are from AI data distribution p′p^{\prime} which is a version of pp with its tail chopped off at rank kk, i.e such that pi′∝pi1[i≤k]p^{\prime}_{i}\propto p_{i}1[i\leq k]. Thus the dataset is drawn from the distribution given by qi=πpi+(1−π)pi′q_{i}=\pi p_{i}+(1-\pi)p^{\prime}_{i}. Test error of a Hutter LLM then writes

Now, thanks to Lemma C.1, it is clear that for any integers 1≤r<R≤∞1\leq r<R\leq\infty and large zz, one has

where c=1−1/β∈(0,1)c=1-1/\beta\in(0,1) and Γ\Gamma is the (upper) incomplete gamma function. Applying (27) with (r,k,z)=(1,k,T)(r,k,z)=(1,k,T) gives

On the other hand, applying (27) with (r,k,z)=(k+1,∞,πT)(r,k,z)=(k+1,\infty,\pi T) and assuming π=Θ(1)\pi=\Theta(1) gives

Putting things together gives the result. ∎

Recall that Bertrand et al. (2023) also formally study such mixtures for iterative retraining. In their setting, they show the existence of fixed points in the mixture proportion that delineates the region of model collapse. These results are complimentary and not contradictory to ours: they combine mixing, large number of iteration, and data-decay, thus studying a combination of effects (under different theoretical conditions, not focusing on scaling laws) that our preceding theorems address separately.

C.4 Grokking for Tail Narrowing

Consider a sample of size TT of which a proportion π\pi comes from the true distribution p=Zip(β)p=Zip(\beta) and the remainder comes from a version p′=Zip(β′)p^{\prime}=Zip(\beta^{\prime}). We have the following scaling law for the Hutter LLM,

where c:=(β−1)/βc:=(\beta-1)/\beta and c′:=(β−1)/β′c^{\prime}:=(\beta-1)/\beta^{\prime}.

Define T‾:=(π/(1−π))−a\overline{T}:=(\pi/(1-\pi))^{-a}, where a:=s/(1−s)a:=s/(1-s), and s:=β/β′s:=\beta/\beta^{\prime}. Then,

(A) Early-Stage Dynamics. For T≲T‾T\lesssim\overline{T}, it holds that Etest≍((1−π)T)−c′E_{test}\asymp((1-\pi)T)^{-c^{\prime}}. Thus, if β′>β\beta^{\prime}>\beta, the money spent on acquiring some clean data is not amortized!

(B) Later-Stage Dynamics. As soon as T≳T‾T\gtrsim\overline{T}, it holds that Etest≍(πT)−cE_{test}\asymp(\pi T)^{-c}. Similarly, we recover the unpolluted sample-size law scaling T−cT^{-c}. For fixed TT and tunable π\pi, this error rate scales like π−c\pi^{-c}.

Let qq be the mixture of pp and p′p^{\prime}. We prove the result for β′≥β\beta^{\prime}\geq\beta; the case β′≤β\beta^{\prime}\leq\beta is analogous. So, one may write

where we have used the fact that (1−π)i−β′≥πi−β(1-\pi)i^{-\beta^{\prime}}\geq\pi i^{-\beta} iff i≤(π/(1−π))−1/(β′−β)=T‾1/βi\leq(\pi/(1-\pi))^{-1/(\beta^{\prime}-\beta)}=\overline{T}^{1/\beta}. The result then follows from (27).∎

Let us conclude by saying that clean data always helps, since EtestE_{test} is decreasing function of π\pi. Indeed, from (26), the derivative w.r.t π\pi is Etest′(π)=−T∑i≥k+1pi2(1−πpi)T−1≤0E_{test}^{\prime}(\pi)=-T\sum_{i\geq k+1}p_{i}^{2}(1-\pi p_{i})^{T-1}\leq 0.

C.5 An interesting detour: Grokking for Fixed-size AI Dataset.

Now consider the scenario where the AI synthesized dataset has fixed size TAIT_{AI} (e.g a frozen chunk of the web), while the clean dataset size is a scalable parameter TrealT_{real}. Taking T=Treal+TAIT=T_{real}+T_{AI} and π=Treal/T\pi=T_{real}/T, we have the following corollary of Theorem 3.2, which includes Corrolary 3.3.

(A) Early-Stage Dynamics. For Treal≪kβT_{real}\ll k^{\beta}, it holds that

(B) Later-Stage Dynamics. As soon as Treal≥CkβT_{real}\geq Ck^{\beta} (where CC is an absolute constant), it holds that

As mentioned in Section 3, AI synthesized data is helpful in the regime where real data is scarce. Once more of real data becomes available the model grokks for a while and then forgets the AI synthesized data to recover the normal scaling law w.r.t TrealT_{real}. Figure 20 gives an illustration of this phenomenon in various settings.

C.6 Proof of Theorem 3.1

where α:=β−1\alpha:=\beta-1. This is because the normalization constant is ∑i≥Npi=∑i≥Ni−β≍N−α\sum_{i\geq N}p_{i}=\sum_{i\geq N}i^{-\beta}\asymp N^{-\alpha}. Now, mix this distribution with qq with equal weights 1/21/2, to obtain a new distribution

For simplicity, assume N≥k+1N\geq k+1 (otherwise, we have all of pp). Build a ”Hutter” LLM from an iid sample of size TT from this distribution (this is equivalent to mixing TT samples from qq and TT samples from π\pi. Then, it is easy to see that the test error is given by

Thanks to previous computations, we know that for large kk, NN, and TT

The first sum is of order T−c(Γ(c,Tk−β)−Γ(c,T))=O(T−c)T^{-c}\left(\Gamma(c,Tk^{-\beta})-\Gamma(c,T)\right)=O(T^{-c}).

The third sum is of order T−c(Γ(c,0)−Γ(c,TNαN−β))=T−c(Γ(c,0)−Γ(c,TN))≍T−cT^{-c}\left(\Gamma(c,0)-\Gamma(c,TN^{\alpha}N^{-\beta})\right)=T^{-c}\left(\Gamma(c,0)-\Gamma(c,TN)\right)\asymp T^{-c}.

The second sum is of order k−α−N−α=((Nk)α−1)N−αk^{-\alpha}-N^{-\alpha}=((\frac{N}{k})^{\alpha}-1)N^{-\alpha}, where α:=β−1\alpha:=\beta-1.

Appendix D Proofs for the Tailed Bigram Model (Section 4)

As a sanity check, with the framework of Equation (17), let us momentarily consider the non-autoregressive setup where p(⋅∣i)=δyip(\cdot\mid i)=\delta_{y_{i}} for all ii, as in classical Hutter. Then, an easy computation shows that

Now, by construction, qT(yi∣i)=1[i∈DT]q_{T}(y_{i}\mid i)=1[i\in\mathcal{D}_{T}]. Thus,

and we recover the classical Hutter result! Thus, our test metric defined in (17) is pointing in the right direction, conceptually.

D.2 Proof of Theorem 4.1

The proof will be based on the results of (Berend & Kontorovich, 2012). ()

Observe that for any choice of mappings π1,π2,…\pi_{1},\pi_{2},\ldots, we have

We deduce that cT(i):=aT(i)+bT(i)≲nT(i)−cc_{T}(i):=a_{T}(i)+b_{T}(i)\lesssim n_{T}(i)^{-c} for any ii. Importantly, the hidden constants don’t depend on ii. Therefore, thanks to [Lemma 9] (Berend & Kontorovich, 2012), we have

where we have used Jensen’s inequality in (*), since the function x↦x−cx\mapsto x^{-c} is concave.

Lower-Bound.

WLOGA summable series of nonnegative numbers (like in aT(i)a_{T}(i) and bT(i)b_{T}(i)) can be reordered without changing the value. consider the following specific choice of permutations defined by πi(j)=j\pi_{i}(j)=j (i.e doesn’t depend on ii). Then,

Thanks to the definition of EtestE_{test} and [Proposition 5] (Berend & Kontorovich, 2012), we deduce that if β∈(1,2)\beta\in(1,2), then

D.3 Proof of Theorem 4.2

It suffices to replace nT(i)n_{T}(i) in (39) and (40) of the proof of Theorem 4.1 with nT(i)∧kβn_{T}(i)\land k^{\beta}, and use the elementary fact that (nT(i)∧kβ)−c=nT(i)−c∨k−βc≍nT(i)−c+k−βc(n_{T}(i)\land k^{\beta})^{-c}=n_{T}(i)^{-c}\lor k^{-\beta c}\asymp n_{T}(i)^{-c}+k^{-\beta c}. The rest of the proof proceeds as that of Theorem 4.1.

D.4 Extensions

Note that the above setup can be extended to the following

Appendix E Proof and Illustration of Triplet Scaling Law (Theorem 5.1)

For any ii, on average it takes 1/pi1/p_{i} iid samples from pp to see the context ii at least once. The effect of tail-cutting at rank kk is effectively to replace the sample size TT by min⁡(T,Tk)\min(T,T_{k}), where Tk=max⁡{1/pi∣i∈[k]}T_{k}=\max\{1/p_{i}\mid i\in[k]\}. In the case where p=Zipf(β)p=Zipf(\beta), we have Tk=1/pk≍kβT_{k}=1/p_{k}\asymp k^{\beta}. On other hand the model (18) proposed in (Cabannes et al., 2023) on Zipf data, the test error writes

where c:=1−1/β∈(0,1)c:=1-1/\beta\in(0,1) and the exponent cq∈(0,∞)c_{q}\in(0,\infty) depends on β\beta and the algorithm qq used to update the embeddings in the memory matrix MTM_{T} in (18). We deduce that tail-cutting at rank kk changes the test error to

Figure 21 confirms the Triplet Scaling Law.

Appendix F Details and Results from the Autoregressive Bigram model with Perplexity

We showcase experiments in the autoregressive bigram model with perplexity loss. We generate sequences of length 100100. Figures 23, Figure 23 and 25 aim to reproduce the ”paired bigram” Figure 12 in this setting, adding a top p mechanism and a temperature mechanism. Figure 27, Figure 27 and Figure 27 regenerates the setting of Figure 6 with the same top p and temperature.

Appendix G Details and Results on Transformer Arithmetic Experiments

Charton (2023) trains sequence-to-sequence transformers to predict the greatest common divisor (GCD) of two positive integers, encoded as sequences of digits in some base BB. He observes that model predictions are deterministic: for any pair (a,b)(a,b) with GCD kk, the model predicts a single value f(k)f(k). Predictions are correct (i.e. f(k)=kf(k)=k) when the GCD is a product of divisors of the base, or of small primes. In all other case, the model prediction is the largest correct prediction (i.e. ll such that f(l)=lf(l)=l) that divides kk. The list of correct predictions L\mathcal{L} varies with the encoding base BB. For instance, for B=10B=10, after 300 million examples, the model correctly predicts L={1,2,4,5,8,10,16,20,25,40,50,80,100...}\mathcal{L}=\{1,2,4,5,8,10,16,20,25,40,50,80,100...\}, the GCD of 2020 and 3030 will be correctly predicted as 1010, but the GCD of 210210 and 140140 will be incorrectly predicted as 1010 (instead of 7070).

We use these models to generate “dirty” training data D(B)\mathcal{D}(B): uniformly sampled pairs of integers (a,b)(a,b) and their (sometimes incorrect) pseudo-GCD, as generated by a trained transformer using base BB. Note: this dataset can be as large as we want. We also create a correct training dataset C(B)\mathcal{C}(B), by sampling pairs (a,b)(a,b) and their correct GCD.

In these experiments, we train models on D(B)\mathcal{D}(B) and C(B)\mathcal{C}(B), for different values of BB. Our goal is to determine whether extensive training on “dirty” data impacts model accuracy.

We focus on 66 bases: B=10,420,1000,2017,2023B=10,420,1000,2017,2023 and 49134913, after training transformers (on correct GCD) over about 300300 millions pairs of integers between one and one million, we achieve the performances listed in Table 1. There, accuracy stands for the proportion of random uniform pairs (a,b)(a,b) that the model can predict correctly, correct GCD is the number of GCD under 100100 that the model correctly predicts (i.e. kk such that f(k)=kf(k)=k), and correct model predictions are the products of numbers in the associated sets. These models are used to generate D(B)\mathcal{D}(B).

In these experiments, all models have four layers, 512 dimensions and 8 attention heads. We consider two architectures: an encoder-only model (17.2M parameters), and an encoder-decoder model (38.7M parameters). The encoder-only model has 2.252.25 times less parameters, trains twice as fast, and incurs no performance penalty.

We then train new models (with the same architecture) to predict GCD, from AI data (generated by the above model), and compare to training with correct data – from correct computation of the GCD. When trained on small number of examples (less than 100100 million), models learning from AI data achieve better accuracy (Table 2). We believe this is due to the fact that AI data smoothes away all the hard case, therefore presenting the model with a cleaner signal in the initial stages.

This pattern changes after extensive training. Table 3 compares performance of models trained on 300M300M and 11 billion examples. For all bases BB, models trained on C(B)\mathcal{C}(B) learn new GCD as training proceeds, whereas models learned on D(B)\mathcal{D}(B) never learn beyond their original performance.

Figures 28 and 29 show that we get the picture predicted by theory: the dirty model learns (until about 300M examples) and then stops learning (while the clean model continues) - its scaling law tapers off as predicted in Theorem 2.1. All the skills the clean model learns after this point are skills the model trained on synthesized data cannot learn (see Figure 30 showing when new learned groups of GCD emerge, and Figure 31 for the learning curve of two models, one trained on the original data, the other on AI data).

We now proceed to train our model on randomly mixed clean and synthesized data for various mixture rates. We train with mixtures of clean and dirty data for mixture fractions of 9%,27%,50%9\%,27\%,50\% and 73%73\% of AI-generated data, for bases 1000, 2023 and 4913, to see the grokking effect. Figure 32 illustrates the results. We can see that even for the average curves over the 10 seeds one can discern a grokking-like delayed learning for the mixtures with relatively small amounts of AI data. This effect can be studied

The models used to generate the data were trained on about 300M examples, and correctly predict 22, 16 and 17 GCD below 100 for bases 1000, 2023 and 4913 respectively. We know (Table 3) that more training on AI-data data only will not improve those performances. On the other hand, we know that models trained on clean data will achieve larger performance. Specifically, out of 10 models trained on clean data, for base 1000, all 10 predict 23 GCD or more after 1.4B examples. The median number of examples needed for the models to predict 23 GCD or more is 465M. For base 2023, 7 models out of 10 predict 17 GCD or more after 2.1B examples. The median number of training samples after which the model bests a model trained on dirty data only is 530M. Finally, for base 4913, 9 clean models out of 10 predict more than 18 GCD after 1.7B examples. The median number of samples is 1.1B.

When zooming in to when the mixture models learn to predict GCD that are ”unlearnable” with an AI-trained model, the grokking effect becomes more apparent.

Table 4 summarizes by listing the time (# of samples) when the mixture models finally learn a GCD that a purely AI-trained model cannot learn, and the delay (in millions samples) since the previous GCD was learned (see also Figure 30 to illustrate the comparison between the clean and the AI-trained model):

The delay period increases with increasing fraction of AI data in the mix. Thus, Table 4 clearly demonstrates the grokking effect of increasing plateau length with fraction of AI data, as predicted by our theoryWe were constrained to stop the experiments at after about 3B samples for most, due to heavy use of compute resources. This probably explains why for the larger AI-mixtures only a few experiments could successfully find new GCDs - the other experiments where still in the pre-grokking phase when they were stopped..

Appendix H Details of Experiments with Llama2

In the realm of large language models (LLMs), the prevailing approach involves a pretraining and finetuning paradigm. For instance, GPT-3 undergoes pretraining on approximately 45TB of text data from diverse sources. This extensive pretraining endows it with a robust capability for a variety of downstream tasks, employing methods such as zero-shot learning, few-shot learning, or finetuning. Our study evaluates the phenomenon of model collapse in scenarios close to the contemporary ‘synthetic data age.’

Utilizing one of the most advanced open-source models, Llama-2 7B, our research investigates the effects on LLMs when they undergo finetuningQuoting (Shumailov et al., 2023), we state that one can, in principle, replicate an experiment described here with training an LLM from scratch to demonstrate scaling law decay. Given that training a single moderately large model produces twice the American lifetime worth of CO2 (Strubell et al., 2019), we opted to not run such an experiment and instead focus on a more feasible finetuning setting. Note that just the language experiments described in the paper took weeks to run. with data generated by other LLMs. To ensure the generation of high-quality data and to provide a relevant but not trivial downstream task, we employ the Wikitext-103 dataset. We segment this dataset into chunks of 128 tokens, between each with a stride of 64 tokens, resulting in approximately 2.2 million chunks. Denote this dataset as D0\mathcal{D}_{0}. The task for generation involves producing the final 32 tokens given the initial 96 tokens from each chunk in the original dataset. In the initial generation (0-th generation), we use the Llama-2 7B FT model, which has been finetuned on D0\mathcal{D}_{0}, applying a generation loss that focuses solely on the cross-entropy loss of the final 32 tokens. We denote this initial model as M0\mathcal{M}_{0}, which demonstrates enhanced capacity for the generation task compared to the standard Llama-2 7B model. By querying M0\mathcal{M}_{0} with the original 96 tokens from D0\mathcal{D}_{0}, we generate the dataset D1\mathcal{D}_{1} and subsequently finetune Llama-2 7B on this dataset to obtain M1\mathcal{M}_{1}. This process is sequentially repeated to generate Di\mathcal{D}_{i} from Mi−1\mathcal{M}_{i-1} and obtain Mi\mathcal{M}_{i} through finetuning. By comparing the performance of various M\mathcal{M} models on the test set derived from Wikitext-103, also segmented into 128-token chunks, we aim to investigate the model collapse in LLMs.

To prevent information leakage across chunks, we restrict the training to only include the loss on the final 32 tokens for all generations. Consequently, the models are never trained on the first 96 tokens coming from the original corpus. The size of the 2.2 million chunks can provide sufficient data for finetuning while avoiding overfitting, given the capacity of Llama-2 7B. Throughout the finetuning process, we maintain consistent settings using learning rate 5e−55e^{-5} for LoRA, using Adam optimizer, dropout rate 0.1, trainable parameter fraction 0.062%. To eliminate the possibility of model collapse due to insufficient sampling and to gain insights into scenarios where more AI-generated data is produced than the model has been trained (or finetuned) on, we consistently utilize a model trained on half the dataset for generating subsequent datasets.

For completeness, we include Figure 34 with loss on the full chunks and Figure 34 that mix the generated data with original data. The mixing curve also aligns well with the grokking phenomenon predicted by theory.

Appendix I More Studies on Tail Cutting and Tail Narrowing Effects

Here, we illustrate how tail cutting in the next-token distribution can lead to tail-narrowing for metrics that take the entire sequence into account, like perplexity. Figure 35 illustrates this for the autoregressive bigram model. This effect is likely due to the combinatorial factors we obtain when considering an additive (or multiplicative) measure like perplexity.