Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language Models

Kushal Tirumala, Aram H. Markosyan, Luke Zettlemoyer, Armen Aghajanyan

Introduction

The rate and extent to which a model memorizes its training data are key statistics that provide evidence about how it is likely to generalize to new test instances. Classical frameworks, such as bias-variance tradeoff , argued for fitting a training set without full memorization. However, recent work has established a more symbiotic relationship between memorization and generalization in deep learning . This paper empirically studies memorization in causal and masked language modeling, across model sizes and throughout the training process.

Much of the recent performance gains for language models have come from scale, with the most recent models reaching up to 101110^{11} parameters . Larger models are also known to memorize more training data , which is a crucial component of their improved generalization. However, perhaps surprisingly, relatively little work has been done in understanding the impact of scale on the dynamics of language model memorization over training. Existing work focuses on analyzing memorization post-training . In this work, we study the memorization and forgetting dynamics in language models, with a focus on better measuring how they change as we scale up model size. Our primary contributions include:

We measure the dependence of memorization dynamics over training on model size (and other factors such as dataset size, overfitting, and learning rate). We find that larger language models memorize training data faster (§ 4).

We design controlled experiments that allow us to characterize the forgetting curves in language models (i.e., how language models naturally forget memories throughout training). Our empirical studies show that forgetting curves have lower bounds — we coin this as the forgetting baseline — and that this baseline increases with model scale, i.e., increasing model scale mitigates forgetting (§ 5).

We analyze the rates of memorization of different parts of speech, finding that nouns and numbers are memorized much more quickly than other parts of speech (§ 4.4). We hypothesize this is because the set of nouns and numbers can be seen as a unique identifier for a particular sample. We provide evidence to this hypothesis by analyzing the rates of memorization in the setting of an existing unique identifier (§ 4.3).

Together, these findings present another piece of the broader puzzle of trying to understand the unique training dynamics that emerge as models grow in size.

Background and Related Work

Memorization in Language Models: Unintended memorization is a known challenge for language models , which makes them open to extraction attacks and membership inference attacks , although there has been work on mitigating these vulnerabilities . Recent work has argued that memorization is not exclusively harmful, and can be crucial for certain types of generalization (e.g., on QA tasks) , while also allowing the models to encode significant amounts of world or factual knowledge . There is also a growing body of work analyzing fundamental properties of memorization in language models . Most related to our work Carlini et al. analyzes memorization of fully trained language models and observes a dependence on model scale, training data duplication, and prompting context length. While we also study scaling behavior, our focus instead is on the memorization dynamics throughout training.

Language Model Training Dynamics: Previous work has extensively analyzed training dynamics to understand how neural models acquire information over training . Saphra and Lopez were the first to analyze training dynamics for language modeling, focusing on the evolution of internal representations over pre-training. This inspired a line of work analyzing how neural language models learn linguistic structure/world knowledge , individual words , and cross-lingual structure over pre-training. This analysis has been extended to many downstream tasks, including text summarization , machine/speech translation , and various NLP tasks .

Forgetting in Language Models: There has also been work studying memory degradation (forgetting) in language models. Catastrophic forgetting or catastrophic interference, first reported in , studies how neural networks tend to forget the information from previous trained tasks or training batches, when trained on new data. This provides a key challenge for continual learning (or life-long learning) , where the goal is to gradually learn from a single pass over a, typically very large, stream of data. A number of mechanisms have been proposed for increasing robustness against catastrophic forgetting . There is also a growing body of work demonstrating that both model and dataset scale can make models more resistant to forgetting , as well as work characterizing how forgetting naturally occurs in image classifiers and how forgetting can improve training efficiency . Machine unlearning is a technique that forces a trained model to forget a previously learned sample , which is primarily motivated by data protection and privacy regulations . Our work is unique in its focus on measuring forgetting during training, and quantifying how it varies with scale.

Scaling Laws: We have consistently seen performance gains by scaling model size , and scale itself has been known to push internal model behavior away from classical bias-variance regimes . Recent efforts have focused on trying to model the scaling laws for language models, including data and model size , applications to transfer learning , routing networks , and various autoregressive generative tasks . While the bulk of work in scaling laws has been empirical, an interesting line of work focuses on theoretically explaining neural scaling laws . Most scaling laws focus only on cross-entropy loss, while we study memorization (defined in § 3).

Experimental Setup

In order to perform a large-scale study of the dynamics of memorization over training, our memorization metric must be reasonably easy to compute but also precise enough to tell us how much the model will actually remember from the training data. Label memorization Label memorization in these prior works usually refers to perfectly fitting a given set of labels is an ideal candidate, because it has consistently provided theoretical insight into underlying properties of neural networks, remains applicable in empirical settings, and is relatively cheap to compute. We formulate our metric as an analog of label memorization for self-supervised settings.

Note that a single word can appear as the ground-truth token for multiple contexts. For a given set of contexts CC (i.e a given training dataset), we can then analyze the proportion of memorized contexts

We refer to this as exact memorization, although it can also be seen as accuracy since we measure how often the argmax of the language model matches the ground truth token. Throughout this work, when we refer to memorization, we will be referring to Definition 1 unless we specify otherwise.

We define τ\tau to be a threshold value for M(f)M(f), and denote T(N,τ)T(N,\tau) as the minimal number of times a language model ff with NN parameter needs to see each training datapoint in order to satisfy M(f)≥τM(f)\geq\tau. When leveraging bigger datasets, models are unable to train for multiple epochs, so we instead consider memorization on a per-update basis. We introduce Mupdate(f,U)M_{update}(f,U) as the memorization on the batch of data on which the model performs the UU’th gradient descent update, and define Tupdate(N,τ)T_{update}(N,\tau) as the minimal number of gradient descent updates a language model with NN parameters needs to perform, to satisfy Mupdate(f,U)≥τM_{update}(f,U)\geq\tau.

Previous work analyzing language modeling memorization defines memorization differently. Motivated by privacy concerns, both and define memorization from a training data extraction standpoint, in which a string ss is extractable if it can be produced by interacting with the language model. More specifically, defines a string ss as being kk-eidetic memorized if it is extractable and appears in at most kk training examples. defines a string ss as kk-memorized if the language model can produce it via prompting with kk tokens of context from training data. This definition only works for causal language modeling because of the dependence on prompting with training data; for masked language modeling uses Definition 1 above. Note that if an example is exactly memorized, it is extractable by definition. In other words, both the set of kk-eidetic memorized tokens and the set of kk-memorized tokens contain the set of exactly memorized tokens (formally, different exactly memorized tokens may be contained in different sets, depending on kk). Therefore, analyzing exact memorization gives a type of lower bound on the kk-eidetic memorization and kk-memorization. In a different line of work motivated by estimating the influence of individual training examples, defines a training example xx as memorized if the difference in expected model performance (where model performance is defined as M(f)M(f) above) over subsets of data including xx and subsets of data not including xx, is sufficiently large. This definition pulls from previous work in theoretically analyzing label memorization in classification settings .

Model Architectures: We replicate publicly available references for Transformer language model architectures . We use the 125M, 355M, 1.3B, 2.7B, 6.7B, and 13B model configurations (see § A.4 for more architectural and training details). We study both causal and masked language models. We train using the FairSeq framework with PyTorch as the underlying framework. For our larger models, we use the fully sharded data-parallel implementation available in FairScale and use Aim experiment tracking .

Datasets: We use two existing datasets across all our experiments: the Wikitext-103 benchmark containing around 103 million tokens , and the RoBERTa corpus used to train the original RoBERTa model, containing around 39 billion tokens (we refer to this as the roberta dataset). We use both datasets in section 4, and primarily use Wikitext-103 in other sections due to computational restrictions.

Larger Language Models Memorize Faster

Larger neural language models are known to be more sample efficient and require fewer optimization steps to reach the same performance while also converging faster , where performance is usually defined as test perplexity. In this section, we study T(N,τ)T(N,\tau) on the training set as a function of NN to answer this question.

In the left plot of Figure 1, we fix a memorization threshold τ=0.9\tau=0.9 and examine T(N,τ)T(N,\tau) as we increase NN. The larger language models need to see each training datapoint fewer times to achieve 90%90\% exact memorization of the training set; in other words, T(N,0.9)T(N,0.9) is monotonically decreasing in NN. When we vary τ\tau between 0.40.4 and 0.950.95 in the right plot of Figure 1, we still observe that T(N,τ)T(N,\tau) is generally decreasing with NN.We fix 0.40.4 as the lower bound for the range because any lower value for the memorization threshold is achieved within the first few epochs across all model scales (the line in Figure 1 is essentially flat), and 0.950.95 as the upper bound because higher values require unreasonably long training time for smaller models. For fixed NN, T(N,τ)T(N,\tau) is increasing in τ\tau, which is expected since memorizing more of the training set requires training the model for more epochs. More interestingly, increasing τ\tau smoothly transitions T(N,τ)T(N,\tau) from constant in NN, to exponentially decreasing in NN (the axes are on a log-log scale).

To investigate the dependence of our observations on the particular language modeling task, we repeat this analysis for the masked language modeling task on WikiText103 with mask probability 0.150.15. Unlike in causal language modeling, Figure 2 shows that T(N,τ)T(N,\tau) is not monotonically decreasing in NN for lower values of τ\tau, and is monotonically decreasing in NN for higher values of τ\tau, where the phase transition”Phase transition” is used in physics to describe significant changes in system behavior that occurs due to varying a parameter, such as temperature. In this case, the parameter is τ\tau between these two regimes occurs between τ=0.6\tau=0.6 and τ=0.7\tau=0.7. Smaller models memorize the training data quicker initially and slower in the long run (e.g., right plot of Figure 11).

Language model training is heavily dependent on the dataset size , and therefore we expect M(f)M(f) to be similarly impacted. In Figure 3, we analyze training set memorization on the much bigger roberta dataset for both masked and causal language modeling. With large datasets such as roberta dataset, it becomes infeasible to perform multiple epochs and evaluate memorization on the entire training set, especially when training larger models. Consequently, we focus on smaller values of τ\tau and investigate the number of gradient descent updates it takes to reach memorization thresholds, i.e., Tupdate(N,τ)T_{update}(N,\tau). In Figure 3 we observe a similar trend as Figure 1, where Tupdate(N,τ)T_{update}(N,\tau) is monotonically decreasing with NN for various τ\tau, in both masked and causal language modeling. Unlike with WikiText103, masked language modeling does not have a phase transition for τ\tau.

2 Why Do Larger Models Memorize Faster?

A natural question at this point is to ask why larger models memorize faster? Typically, memorization is associated with overfitting, which offers a potentially simple explanation. In order to disentangle memorization from overfitting, we examine memorization before overfitting occurs, where we define overfitting occurring as the first epoch when the perplexity of the language model on a validation set increases. Surprisingly, we see in Figure 4 that as we increase the number of parameters, memorization before overfitting generally increases, indicating that overfitting by itself cannot completely explain the properties of memorization dynamics as model scale increases.

The learning rate is not constant across our training configurations. Intuitively, larger learning rates should lead to quicker memorization. To investigate to what extent our results can be explained by learning rate, we take a subset of the architectures available above and train on the WikiText103 dataset across a standard range of learning rates while measuring memorization, in Figure 5. Even if we fix a learning rate, larger models reach 0.90.9 memorization faster, suggesting that our results are not caused solely by differences in learning rates. Interestingly, sensitivity to learning rate generally decreases as we increase the model size. We also notice in Figure 5 that T(N,τ)T(N,\tau) goes down initially (for low LRs) and eventually rises (for high LRs), and as the long as the chosen learning rate places us near the lowest point on the curve, the memorization dynamics do not change significantly (note that axes are on log-scale). This result is consistent with the growing intuition that for neural language models past a particular scale, the learning rate is not a significant hyperparameter .

Exhaustively searching all such possible factors is intractable, and providing a complete explanation for why larger models memorize faster is outside the scope of this work. Instead, in the following sections, we present studies that we hope will expand the toolkit for answering such questions.

3 Memorization via. Unique Identifiers

Recent work studies how to use external memory to improve performance . In this subsection, we question whether such architecture changes are necessary. Motivated by information retrieval systems, we take a simple approach — we prepend a unique identifier to every example in the training set and examine whether memorization speed increases. Specifically, we fix the language modeling task as causal language modeling on WikiText103 with the 125M parameter model, and in front of every training example, we insert the string document ID where unique_id is a unique integer, one for each training context. In order to utilize all these unique integers, we must add them to the dictionary of tokens, which causes a significant increase in the model size since the last layer in the language model must have an output dimension equal to the size of the dictionary. Therefore, any change in M(f)M(f) dynamics could be attributed to the extra parameters we add from increasing dictionary size. To control for this, we first examine the effect of just increasing dictionary size (without using any of the added tokens). Then, we utilize those added tokens to prepend every training example and observe the change in M(f)M(f) dynamics. In Figure 6, we see that increasing the dictionary size does improve the speed of memorization. Even though we previously demonstrated that larger models memorize faster, this is still surprising considering that we do not increase parameter size in a significant way — we are effectively adding fake tokens to the dictionary. Moreover, when we leverage those added tokens to identify training examples uniquely, we see yet another gain in memorization, although prompting using a document ID shifts memorization dynamics away from being monotonically increasing over time.

4 Memorization Through the Lens of Parts of Speech

In the previous section, we showed that a unique identifier enhances memorization. Regular text also contains strong proxies to unique identifiers in the form of numerals and proper nouns. Motivated by this, we study syntactic features of memories using part-of-speech (POS) tagging.We use spaCy to identify parts of speech in a text. We track the ratio R(p)R(p) of the number of positions for which the part of speech pp was correctly predicted to the total number of tokens in the ground truth tagged with that part-of-speech pp (left plot in Figure 7). In the right plot of Figure 7 we show a similar ratio, denoted Rmem(p)R_{mem}(p), but the numerator only considers the tokens that are also exactly memorized. The correctly predicted part of speech does not necessarily imply exact memorization, which is clearly illustrated by Figure 7 where we see the language model memorizing parts of speech faster than the exact value of the token. While all parts of speech are eventually memorized, some parts of speech are memorized faster, which aligns with previous work . However, unlike previous workThis difference could be due to model family (we use causal LMs while previous work uses masked LMs), we find that nouns, proper nouns, and numerals are memorized noticeably faster than verbs and adjectives, both in terms of R(p)R(p) and Rmem(p)R_{mem}(p). This has potential implications for privacy, since sensitive information is likely to be a noun/proper noun/numeral. Our findings also very loosely align with work studying child language acquisition .

Forgetting Curves in Language Models

This section studies the dual of memorization — forgetting in language models. Inspired by the forgetting curve hypothesis, according to which human memory declines over time when there is no attempt to retain it , we are interested in understanding the dynamics of memory degradation in language models.

We first choose a batch of data not available in the training set, i.e. a batch of data from a validation set. We refer to this batch of data as the special batch. We then take a checkpoint from model training, plug in the special batch so that the model can train on it, and resume standard training on the training set. We then evaluate how memorization degrades on the special batch and analyze the various factors the forgetting curve may depend on. We use the entire validation set as the special batch throughout this section. The special batch is only seen once when it is immediately introduced.This experimental setup is different from catastrophic forgetting, as we fix the data distribution by pulling the special batch from the same dataset as the training set. Similarly, it differs from machine unlearning since we are not algorithmically removing information from a language model; instead, we analyze natural forgetting. It is also different from intrinsic hallucination , where there is an assumption that contradicted output is semantically correct (e.g., the language model outputs a wrong date).

In the left plot of Figure 8, we show the forgetting curve for the 2.7B model. Exact memorization on the special batch degrades quickly at first, but slows down exponentially as we continue training The average sequential difference in memorization (on the special batch) on the last 33 epochs of training is at most on the order of 10−310^{-3}, whereas the average sequential difference in the first 33 epochs of training is consistently on the order of 10−210^{-2} (see Figure 15 in § A.2.2). In other words, the forgetting curve on the special batch seems to approach a baseline — we refer to this trend as the forgetting baseline. We approximate the forgetting baseline by looking at the lowest memorization value on the special batch throughout training.

We show the forgetting baseline as a function of the model scale in the right plot of Figure 8. We see that the numerical value for the baseline is monotonically increasing with the model scale. This implies that larger models forget less, aligning with recent work studying catastrophic forgetting on image classification tasks . This is beneficial because larger models can leverage more information from previous tasks; however, from a privacy perspective, this is not ideal because it implies larger models may be potentially retaining more sensitive information from training data.

We also investigate the sensitivity of the forgetting baseline on data batch order. In Figure 9, we perform the same forgetting curve analysis described above but start the analysis at different training checkpoints (we start at the 14th, 39th, and 63rd epochs). This way, we alter the order of the data batches given to the model (since the special batch will appear in a different place in the global order of data batches given to the model) without drastically changing the experimental setup. We observe that the forgetting baseline is not sensitive to data batch orderThe max difference between the numerical values for the baseline are on the order of 10−310^{-3}.

Motivated by replay methods from continual learning (see for a survey) and work in promoting retention memories through repetition in both humans and neural models , in Figure 10 we study the effect of repetition (left) and spaced repetition (right) on the forgetting baseline. In the left plot, we inject the special batch into the training set multiple times before continuing training on the training set alone. We observe that the forgetting baseline is monotonically increasing as a function of repetition frequency (differences in the baseline value are on the order of 10−210^{-2}). To study the spaced repetition, we periodically inject the held-out set into the training set, train on it once, and then continue training on the training set alone. We see in the right plot of Figure 10 that spaced repetition incurs minimal effect on the forgetting baseline (on the order of 10−310^{-3}), independent of the length of spacing between the repetitions.

An exciting direction for future work will be to understand the structure of the baseline — for example, understanding what types of tokens (parts of speech, synonyms, facts, syntax) are memorized in the baseline and the overlap of tokens memorized in the baseline with tokens in the training set.

Conclusions and Discussion

We study the properties of memorization dynamics over language model training and demonstrate that larger models memorize faster. We also measure the properties of forgetting curves and surprisingly find that forgetting reaches a baseline, which again increases with the model scale. Combined with memorization analyses that expose the unintuitive behavior of language models, we hope to motivate considering memorization as a critical metric when increasing language model scale.

Most work studying memorization in language modeling is primarily motivated by privacy (see § 2). While theoretically, there are well-established frameworks to quantify privacy such as differential privacy , empirical privacy in language modeling is not well-defined — does memorizing common knowledge count as information leakage? Does outputting a synonym count as harmful memorization? As per our Definition 1, we implicitly focus on information that is sensitive if outputted verbatim (phone numbers, SSNs, addresses, medical diagnoses, etc.), rather than capturing all aspects of privacy. It is also known that text data used for training language models contain certain biases and stereotypes (e.g., ); therefore, our work has similar implications for how long language models can train before they definitively memorize these biases from training data.

We also hope our work highlights the importance of analyzing memorization dynamics as we scale up language models, instead of only reporting cross entropy. Cross-entropy loss and memorization capture different behavior — for example, in many of our memory degradation experiments, even though memorization approaches a baseline, we observe that perplexity is still increasing (see Figure 14 in § A.2 for an example). This implies that the model is becoming unconfident about its exact predictions, which we can only conclude because we inspect both loss and memorization. More importantly, the forgetting baseline behavior would be entirely obscured if we did not inspect memorization dynamics. Similarly, there are multiple instances where we uncover interesting behavior because we focus on memorization dynamics (§ 4.4, § 4.3, § A.3), rather than focusing only on cross-entropy loss.

Acknowledgements

The authors would like to thank Adina Williams, Chuan Guo, Alex Sablayrolles, and Pierre Stock, for helpful discussions throughout the course of this project. The authors would also like to researchers at FAIR who commented on or otherwise supported this project, including Shashank Shekhar, Candace Ross, Rebecca Qian, Dieuwke Hupkes, and Gargi Ghosh.

References

Appendix A Appendix

For completeness, in this section we plot our memorization metric M(f)M(f) over training for all model sizes. In any of these plots, observe that taking a horizontal slice for a fixed τ\tau is equivalent to computing T(N,τ)T(N,\tau). In Figure 11, we plot M(f)M(f) over training for WikiText103. We see that generally (across language modeling tasks and and values of τ\tau), larger models memorize faster. We do notice a caveat in Figure 11, where we observe that in initial stages of training, smaller models memorize faster, but larger models eventually surpass smaller models.

When we analyze larger datasets, performing multiple epochs of training becomes infeasible, and so we track memorization with each gradient descent update. Similarly, we cannot analyze M(f)M(f) for the entire training dataset. We use notation introduce in § 1, specifically Mupdate(f,U)M_{update}(f,U) where UU is the number of gradient updates performed on model ff. This quantity is defined as the memorization on the batch of data given to the model on the UU’th update. In figure 12, we take a rolling average with window size 55 when plotting Mupdate(f,U)M_{update}(f,U) to smooth out curves.

To check that Mupdate(f,U)M_{update}(f,U) is a viable proxy for M(f)M(f), in Figure 13, we plot both M(f)M(f) and Mupdate(f,U)M_{update}(f,U) up to 30000 updates for two model sizes. We fix 3000030000 as the upper bound, because we only train some model sizes up to 3000030000 updates in the roberta experiments in § 4, and therefore can only completely assess the impact of scale on Mupdate(f,U)M_{update}(f,U) dynamics up to 3000030000 updates. We see that Mupdate(f,U)M_{update}(f,U) has periodic behavior, but overall does not deviate too much from M(f)M(f).

We note that Definition 1 is not the best way to study memorization: it ignores model confidence and it does not normalize for duplication in the training set (it is known that duplication in the training set helps models memorize tokens ). However, as mentioned in Section 3, all previous definitions of memorization seem to involve Definition 1 in some form. In this way, we study a metric fundamental to memorization regardless of the precise definition of memorization.

A.2 Forgetting Baseline Analysis

This section shows how perplexity and memorization on the special batch evolve over training. In Figure 14 we see that perplexity continues to increase over training, while memorization flatlines. This is a clear experimental setup where we find cross-entropy loss capturing different behavior from memorization. We show plots for the 1.3B model scale, although all of the experiments in § 5 exhibit very similar trends.

A.2.2 Verifying Existence of Baseline

To verify the existence of the forgetting baseline discussed in § 5, we observe the sequential difference in M(f)M(f) of the special batch, from epoch to epoch. More formally, if M(f)TM(f)_{T} denotes the memorization at epoch TT, we investigate diff(T)=M(f)T−M(f)T−1\texttt{diff}(T)=M(f)_{T}-M(f)_{T-1} on the special batch, for T>1T>1. In Figure 15 we show this plot for a few model scales, and we clearly see that the sequential difference in M(f)M(f) exponentially approaches .

A.3 Analyzing Memory Unit Length Over Training

This section investigates a fundamental property of memories — memory unit length LL. We look at individual tokens memorized as having length L=1L=1, memorized bigrams as having length L=2L=2, memorized trigrams as having length L=3L=3, etc. Analyzing memory length is interesting because it has implications for how language models retain nn-grams, which are an important part of language. Moreover, recent work shows that chain-of-thought prompting improves language model performance ; understanding memory unit length informs us whether a similar method might work for improving performance when training (if a language model has low memory unit length, then including chain-of-thought-type texts in the training set might not have a significant effect). An empirical side note is that these experiments were run separately from the main paper experiments, so we provide original M(f)M(f) curves for reference.

We track the average value of LL across the entire training dataset for causal language modeling on WikiText103. Note that in our all our experiments, the sequence length is constrained to be less than 512512 tokens, with an average sequence length of 430.12430.12 on WikiText103. In the left plot of Figure 16 we analyze the average memory unit length over training for two model sizes. We observe across model sizes that average memory unit length steadily increases over time, roughly taking a sigmoidal shape. We notice that the larger 2.7B model has an average LL increasing faster than the 125M model. This is consistent with our previous results because we know larger models memorize, and some of these tokens are likely to be adjacent to each other, especially as the model achieves higher values of M(f)M(f). Surprisingly, we see that the average memory unit length is much lower than the average sequence length of 430.12430.12, suggesting that even with high individual token memorization (which is achieved as shown in the right plot of Figure 16), there are always tokens in the middle of a text that the language model has not yet memorized, which break up the memories.

A.4 Model Training/Dataset Details

In this section, we layout the details of experiments, although most training details we pull directly from publicly available references . As such, we provide the details of model architectures using the same style as Table 1 in for ease of comparison. All models use GELU activation for nonlinearity. We leverage the Adam optimizer , with β1=0.9\beta_{1}=0.9, β2=0.98\beta_{2}=0.98, and ϵ=10−8\epsilon=10^{-8}. For reproducibility, we set weight decay to , dropout to , and attention dropout to . We use a polynomial learning rate schedule, and following we scale up our learning rate from to the maximum learning rate over 375M375M tokens, and scale down to over the remaining T−375MT-375M tokens (for all masked language modeling experiments, and all roberta experiments, we have T=300BT=300B; for causal language modeling experiments on WikiText103 we have T=100BT=100B). We fix a sequence length of 512512 across all experiments, but we break input text up into complete sentences, so not all input texts have length exactly equal to 512512. In masked language modeling experiments, we use a mask probability of 0.150.15. When training language models, we use the standard procedure of minimizing cross-entropy loss, and use dynamic loss scaling .

As mentioned in § 3, we use FairSeq which relies on PyTorch . When training models, we leverage fully sharded data-parallel implementation of models in FairScale . We utilize NVIDIA A100 GPUs with 40GB of memory. Increasing model scale requires different amounts of GPUS: 125M and 355M generally required 16 GPUS, 1.3B required 32 GPUS, and 2.7B, 6.7B, and 13B generally required 64 GPUS (although some experiment runs were launched with 128 GPUS in order to decrease training time). Exact training time varied depended on model scale and dataset size, but all models were trained for up to 140 hours.

In both datasets we use, there is a possibility for sensitive or offensive text to be included in the training set, since both benchmarks use data that is from the Internet. We also note that the WikiText103 benchmark we use throughout the work is available under the Creative Commons Attribution-ShareAlike License. The roberta dataset we use refers to the corpora of text originally used to train the RoBERTa model (see ). This dataset not publicly available under any license, however subsets of data that make up the corpus are publicly available.