Proving Test Set Contamination in Black Box Language Models
Yonatan Oren, Nicole Meister, Niladri Chatterji, Faisal Ladhak, Tatsunori B. Hashimoto
Introduction
Large language models (LLMs) have driven remarkable improvements on a number of natural language processing benchmarks (Wang et al., 2019) and professional exams (OpenAI, 2023). These gains are driven by large-scale pretraining on massive datasets collected from the internet. While this paradigm is powerful, the minimal curation involved has led to growing concerns of dataset contamination, where the pretraining dataset contains various evaluation benchmarks. This contamination leads to difficulties in understanding the true performance of language models – such as whether they simply memorize the answers to difficult exam questions. Disentangling the effects of generalization and test set memorization is critical to our understanding of language model performance, but this is becoming increasingly difficult as the pretraining datasets are rarely public for many of the LMs deployed today.
Although there is ongoing work by LLM providers to remove benchmarks from pre-training datasets and perform dataset contamination studies, such filtering can fail due to bugs (Brown et al., 2020a), be limited to a select set of benchmarks (Brown et al., 2020a; Wei et al., 2021; Chowdhery et al., 2022), and requires trust in these vendors. Increasing competitive pressures have also led to some recent model releases to include no contamination studies at all (OpenAI, 2023). These factors make it critical for us to be able to audit existing language models for the presence of benchmark datasets without the cooperation of language model providers.
In parallel to contamination studies, there has been a growing literature on heuristic membership inference algorithms, that seek to reverse engineer aspects of the pretraining dataset (Carlini et al., 2019; Mattern et al., 2023) as well as provide some evidence for test set contamination (Sainz et al., 2023; Golchin & Surdeanu, 2023). However, the heuristic nature of these methods limits their usefulness, as these methods cannot elevate speculation about a suspected instance of test set contamination into an irrefutable proof of contamination.
In this work, we show it is possible to go beyond heuristics and provide provable guarantees of test set contamination in black box language models. More specifically, we provide a statistical test that can identify the presence of a benchmark in the pre-training dataset of a language model with provable false positive rate guarantees and without access to the model’s training data or weights.
To achieve these guarantees, we exploit the fact that many datasets have a property known as exchangeability, where the order of examples in the dataset can be shuffled without affecting its joint distribution. Our key insight is that if a language model shows a preference for any particular ordering of the dataset – such as a canonical ordering that appears in publicly available repositories – this violates exchangeability and can only occur by observing the dataset during training (Figure 1). We leverage this insight to propose a set of tests that compares the language model’s log probability on the ‘canonical’ ordering (taken from public repositories) to the log probability on a dataset with shuffled examples, and flag a dataset if the two log probabilities have statistically significant differences.
Using these ideas, we propose a computationally efficient and statistically powerful test for contamination which shards the dataset into smaller segments and performs a series of log probability comparisons within each shard. We prove that this sharded test provides control over the false positive rate, enables computationally efficient parallel tests, and substantially improves the power of the test for small p-values.
We evaluate our statistical test on a 1.4 billion parameter language model trained on a combination of Wikipedia and a curated set of canary test sets. Our test is sensitive enough to identify test sets with as few as 1000 examples, and sometimes even appearing only twice in the pretraining corpus. In the case of higher duplication counts, such as datasets appearing 10 or more times, we obtain vanishingly small p-values on our test. Finally, we run our test on four commonly used, public language models to study the behavior of our test on language models in the wild and find little evidence of pervasive and strong test set contamination.
Demonstrating the use of exchangability as a way to provably identify test set contamination using only log probability queries.
Construction of an efficient and powerful sharded hypothesis test for test set contamination.
Empirical demonstration of black-box detection of contamination for small datasets that appear few times during pretraining.
Our three contributions suggest that black-box identification of test set contamination is practical and further improvements in the power of the tests may allow us to regularly audit language models in the wild for test set contamination. To encourage the development of new provable guarantees for test set contamination, we release our pretrained models as a benchmark for developing future statistical tests.https://github.com/tatsu-lab/test_set_contamination.
Problem Setting
Our high-level goal is to identify whether the training process of a language model included dataset . In our setting, the only method we have to study is through a log probability query for a sequence (i.e. no access to dataset or parameters). This setting mirrors many common situations with API-based model providers (Brown et al., 2020b; Bai et al., 2022) and matches an increasing trend where the training data is kept secret for ‘open’ models (Touvron et al., 2023; Li et al., 2019).
Provably identifying test set contamination can be viewed as a hypothesis test in which the goal is to distinguish between two hypotheses:
: is independent of
: is dependent on
where we treat as a random variable whose randomness arises from a combination of the draw of the pretraining dataset (potentially including ) and we will propose a hypothesis test with the property that it falsely rejects the null hypothesis with probability at most .
In most cases, we can make use of a property of a dataset known as exchangeability to obtain our false positive guarantee. Nearly all datasets can be expressed as a collection of examples where the ordering of the examples are unimportant, and the probability of any ordering would be equally likely (i.e. for any permutation ). Notably, this assumption would hold under the standard assumption that the dataset is a collection of i.i.d examples.
Whenever exchangability of the dataset holds, the log probabilities of the model under must have a useful invariance property,
Let be a function that takes a dataset and concatenates the examples to produce a sequence, and let be a random permutation of the examples of where is drawn uniformly from the permutation group. For an exchangeable dataset and under ,
Proof This follows directly from the definitions of exchangability and . Since is exchangable, and by the independence of from under , we know that . Thus, the pushforward under must have the same invariance property. ∎
Proposition 1 is the basic building block of our tests. It implies that the log probabilities of under have the same distribution when shuffled, and this permutation invariance will enable us to directly apply standard results on constructing permutation tests (Lehmann & Romano, 2005).
The false positive rate guarantee holds with extremely weak assumptions, but a useful test should also have high power, meaning that it should have a high detection rate under . We cannot hope for high detection rate without further assumptions. For instance, an adversary may hide an encrypted copy of within the parameters of the model (which would induce a clear dependence between the model and ) but it would be nearly impossible for us to detect such a situation even with weight access.
However, most existing forms of contamination are benign. In the benign contamination setting we consider, pretraining datasets become contaminated when test sets accidentally slip through filtering mechanisms Brown et al. (2020a). In this case, we have a reasonable expectation that the invariance in proposition 1 will be violated and as the language model is explicitly trained to maximize the log-likelihood over its training data, including . The violation of exchangability allows us to reliably detect test set contamination, and the existing literature on memorization (Carlini et al., 2021) suggests that many models may verbatim memorize the order of examples in a benchmark dataset. We now focus on building tests that can reliably identify this form of memorization.
Methods
The core idea of our statistical test is to compare the log probability of the dataset under its original ordering to the log probability under random permutations. We begin by describing the basic version of this idea, which directly implements a permutation test on the log probabilities. We then identify some drawbacks of this approach and describe a sharded test which improves the statistical power and computational efficiency of the test.
Under the null hypothesis, the likelihood under the model of any permutation of the dataset has the same distribution, and thus the rank of among any set of randomly permuted probabilities will be a uniform random variable over (Lehmann & Romano, 2005, Theorem 15.2.2).
This can be used directly to construct a permutation test. Consider the proportion of permuted copies of with lower log-likeihood than the canonical ordering under the model,
The distribution of will be uniform under , and we can test for contamination at a significance level by rejecting when . In practice, computing this expectation over all is intractable, and we replace this with a Monte Carlo estimate and the appropriate finite-sample correction (Phipson & Smyth, 2010), which gives
This test is simple and straightforward to implement, and the validity of this test when rejecting at is clear from standard results on permutation testing (Lehmann & Romano, 2005; Phipson & Smyth, 2010). However, this test suffers from a major drawback in its Monte Carlo implementation – the runtime of the test in terms of the number of log probability computations is for a sequence of length and the p-value can never be below . For hypothesis tests that aim to reject at very low p-values (or with substantial multiple hypothesis testing corrections), this poses a tradeoff between statistical power and computational requirements.
2 A sharded likelihood comparison test for contamination
What are some drawbacks of the naive permutation test? It has an undesirable tradeoff between statistical power and computational requirements for small , and also requires that the model assign higher likelihood to the canonical ordering than nearly all shuffled orderings of . This latter condition can also be a serious problem, as the model may have biases the prefer certain orderings (e.g. ones that place duplicate examples next to each other) regardless of the order seen during training.
We show that this is possible and the resulting test resembles a series of log probability comparisons followed by a t-test to aggregate these results. More specifically, we will partition the examples into contiguous shards formed by grouping together adjacent examples
where each shard contains at least examples.
Then, we will permute the examples within each shard and compare the likelihood of the canonical ordering to a Monte Carlo estimate of the average likelihood of the shuffled ordering as
Finally, to construct the test, we aggregate these shard statistics via the mean and test for whether is zero-mean using a t-test.
This statistical test, whose pseudocode is given in Algorithm 1, addresses the shortcoming of the permutation test by converting a single rank comparison into a collection of log probability comparisons. The t-test based approach also requires runtime for permutations, but there is no minimum p-value, and in practice we find that the p-values obtained by this approach decay rapidly, as it only requires that the language models consistently assign higher-than-average log probabilities to the canonical ordering, rather than requiring that the canonical log probability be in the tails of the permutation null distribution.
Under the null, we expect to be the sum of independent random variables and we can now show that the overall test provides a false positive rate guarantee.
Under the null hypothesis, an i.i.d dataset , and finite second moments on ,
as and is defined as the p-value in Algorithm 1.
Proof The result follows directly from the combination of Proposition 1 and standard invariance results in (Lehmann & Romano, 2005). First, by Proposition 1, note that the distribution of is invariant to the permutation .
By Lehmann & Romano (2005, Theorem 15.2.2), this guarantees that the permutation distribution is uniform over the support, and the statistic must be zero-mean. Next, we note that each shard is independent, as each example is split independently into a separate shard with no overlap. By independence and the finite second moment condition, under the null by the central limit theorem and a one sided t-test provides asymptotically valid p-values with uniformly as (Lehmann & Romano, 2005, Theorem 11.4.5). ∎
This result ensures that the sharded rank comparison test also provides (asymptotic) guarantees on false positive rates, much like the permutation test. The test we propose here has two small differences relative to the permutation test – it provides asymptotic, rather than finite-sample valid p-values and assumes i.i.d for the proof. These conditions could be relaxed by the use of Berry-Esseen bounds to obtain finite-sample convergence rates for the CLT as well as replacing our use of a standard central limit theorem with one applicable to the sums of exchangable random variables. However, we opted to present the simpler asymptotic test given the frequent use of i.i.d data generation assumption in the literature as well as the fast convergence of the CLT in practice.
Experiments
We now demonstrate that our test is effective for detecting many common forms of test set contamination. We begin by training a 1.4 billion parameter language model, consisting of both Wikipedia and a known collection of exchangeable test sets. These canaries serve as positive controls for our test, and our goal will be to flag as many of these as possible. Having validated the test in a setting with known contamination, we then explore its use with existing open models.
To validate our test statistic, we train a 1.4 billion parameter GPT-2 model from scratch with a combination of standard pretraining data (Wikitext, taken from the RedPajama corpus (Together Computer, 2023)) and known test sets. We derive 10 test sets from numerous standard datasets (BoolQ (Clark et al., 2019), HellaSwag (Zellers et al., 2019), OpenbookQA (Mihaylov et al., 2018b), MNLI (Williams et al., 2018), Natural Questions (Kwiatkowski et al., 2019a), TruthfulQA (Lin et al., 2022), PIQA (Bisk et al., 2019), MMLU (Hendrycks et al., 2021)), and subsample the datasets to around 1000 examples to ensure that the test sets remain a small part of the overall pretraining dataset (See Table 1 for exact sizes). While we do not know if these datasets are exchangable when they were constructed, we can make them exchangable simply by applying a random shuffle to the dataset, which would make all orderings of the examples equally likely.
To test our ability to detect benchmarks at various duplication rates, we duplicate each of the datasets a different number of times - ranging from 1 to 100 (See Table 1). The overall pretraining dataset has 20.2B tokens, with 20M tokens associated with some benchmark dataset.
Test parameters
The sharded rank comparison test requires two additional parameters: the shard count and the permutation count . Thoughout these experiments we use shards and permutations. In our ablations below, we found that the tests are not particularly sensitive to these parameters, and we fix these parameters to avoid the possibility of p-hacking.
Canary Results
In Table 1, we find that our test is highly sensitive, and provides near-zero p-values at duplication rates of 10 or above. These detections hold for relatively small datasets ( examples) and for a modestly sized language model with 1.4 billion parameters. Given that many test sets are much larger in practice, and many language models of interest are much larger and memorize more aggressively (Carlini et al., 2019), these findings suggest that our test is likely to detect contamination in practice.
While the permutation test attains significance (at a typical , say) for all benchmarks duplicated at least 10 times, the p-values are bounded below by , where the number of permutations used here is . Results for our sharded test use ; even with half the compute, the sharded test attains comparable performance for benchmarks with small duplication rate. However, the p-values attained by the sharded test for moderate to high duplication rates are vanishingly small.
Attaining comparably low p-values using the permutation test is computationally infeasible. For example, to allow for the possibility of a p-value as low as 1.96e-11 (matching the MNLI result) would require permuting the dataset times, and as many forward passes of the model.
Although our test is unable to detect contamination at a duplication rate of 1, other existing literature on memorization has suggested that detection at this duplication level is extremely difficult. Prior work has found that existing tests of memorization begin to work with 10-30 duplicates (Carlini et al., 2021), that deduplicated text is hard to extract (Kandpal et al., 2022), and that dataset contamination with a duplication rate of 1 barely affects downstream benchmark performance (Magar & Schwartz, 2022).
Power as a function of duplication rate.
We carefully study the lowest duplication rate for which our test can reliably detect contamination. To do this, we perform the above canary study but with duplication rates ranging from 1 to 7, and we show the aggregate log p-values for each duplication rate in Figure 2. We find that we cannot reliably detect duplication rates of 1, but that at counts of 2 and 4 we begin to detect some test sets (gray points below the dashed line) and that the detection threshold is around a duplication rate of 4. This suggests that even small amounts of dataset duplication would be sufficient for detection, and future improvements to the power of this test could enable reliable detection at much lower duplication rates.
A public benchmark of provable test set contamination
Our work demonstrates that exploiting exchangability allows us to provably detect test set contamination even for small, low-duplication count datasets. However, it is an open question whether there are tests that can reliably detect contamination at a duplication rate of 1. To support future work on this open problem, we release our pre-trained models trained on Wikitext mixtures together with the corresponding canary test sets.
In addition to the release, we will maintain a leaderboard of methods that provide (asymptotically valid) p-values, ranking methods by the average log p-value. We hope that the model and benchmark spurs further development of tests for contamination, and encourage members of the research community to improve on our results for low duplication counts.
2 Sharding and Permutation Count
Our test relies on two parameters – the number of shards in the test, and the number of permutations to sample. Both of these affect the power of the test, and we carefully study the impact of these parameters on our ability to detect test sets by evaluating our pre-trained model on the 6 datasets that contain 1000 examples (BoolQ, HellaSwag, MNLI, NaturalQuestions, TruthfulQA, PIQA). For the number of shards, we explore a range of settings, from 10 shards to 200 shards and for permutations we test a range from 1 to 50 permutations.
Our results in Figure 3(a) show that there is a sweet spot to the number of shards, around 10-20 shards, where our detection rate for test sets are maximized. Larger numbers of shards perform worse, since each shard involves fewer examples. Shards below 10 do not perform well, as this is likely too few samples to merit the use of an asymptotically valid test like the t-test.
Permutation count sensitivity
We also measure the dependence of our test on the number of permutations per shard in Figure 3(b), and find more permutations to generally improve the power of our test. We test permutations of 1, 2, 10, 25, 50 and compute the average log p-value of the 6 datasets evaluated on the pretrained model. In practice we find that there is substantial diminishing returns beyond 25 permutations in the t-test. This stands in stark contrast to the permutation test, where a permutation count of 25 would only allow for a minimum p-value of 0.038.
3 Evaluating existing models for dataset contamination
We now demonstrate the utility of our procedure in validating test set contamination in multiple publicly available language models: LLaMA2 (Touvron et al. (2023)), Mistral-7B (Mistral (2023)), Pythia-1.4B (Biderman et al. (2023)), and GPT-2 XL (Radford et al. (2018)), on eight public test benchmarks: AI2-Arc (Clark et al. (2018)), BoolQ (Clark et al. (2019)), GSM8K (Cobbe et al. (2021)), LAMBADA (Paperno et al. (2016)), NaturalQA (Kwiatkowski et al. (2019b)), OpenBookQA (Mihaylov et al. (2018a)), PIQA (Bisk et al. (2019)), and MMLU (Hendrycks et al. (2021)). Computationally, we find that our test runs reasonably quickly for a 7 billion parameter model, allowing for the testing of 65 files for contamination in under 24 hours using 50 shards with 250 permutations per shard, and we find that the test outcomes are in general agreement with the contamination study results of Brown et al. (2020c) and Touvron et al. (2023): we do not find evidence of pervasive verbatim test set contamination across the models and benchmarks we tested.
We tested five models for contamination by eight publicly available benchmarks and list the results in Table 2. We use 50 shards and 250 permutations per shard throughout the experiments. For test sets containing more than 5,000 examples, we truncate and test only the first 5,000. Benchmark datasets in the wild may contain non-exchangable elements such as sequential indexes or duplicate examples which would break the false positive guarantees of our test. To check for these possibilities we manually inspected each dataset for non-exchangable elements, and also run our tests on a ‘negative control’ of BioMedLM (Bolton et al., 2022), a language model trained exclusively on PubMed data and known not to contain the benchmarks used here. The p-values computed for BioMedLM are not significant across the benchmarks shown here, suggesting that any significant results for the other models tested are not simply due to non-exchangeability.
Our results in Table 2 show non-significant results across most models and datasets. While failure to reject the null hypothesis is not direct evidence in favor of the null, our synthetic canary evaluations from section 4.1 suggest that it is unlikely for these models to have seen most of these test sets more than 2 to 10 times. One notable exception is AI2-ARC on Mistral, which does return a low p-value of 0.001 and could suggest some contamination. While this p-value appears small, we caution the reader that applying a multiple hypothesis test corection would imply that 0.001 is right at the boundary of statistical significance, and due to challenges with garden-of-forking-paths type analysis, significance tests that are at the boundary of the rejection cutoff should be interpreted cautiously. We present these results as showing promising first steps towards third-party audits of test set contamination, rather than a direct proof of contamination of specific models.
We additionally discuss contamination tests on MMLU, which was identified as a potential contaminant in recent studies (Touvron et al. (2023)), and involves important details. MMLU is not a single test set, but rather a collection of test sets. We speculate that at least 14 of the test sets are non-exchangeable, and applying our test directly would break the false positive guarantees. To understand if our tests can still provide some insights into contamination, we run our MMLU test with a non-exchangability filtering criterion.
To evaluate models for contamination by MMLU, we first exclude those 14 test files from consideration for which our test flags either BioMedLM or GPT-2 as contaminated (both are negative controls as the GPT-2 family of models predates MMLU). We run our test on each of the 43 remaining test files (with 1000 permutations, due to the small size of each file) and aggregate the p-values using Fisher’s method (Fisher (1934)). Although the omnibus p-values resulting from this procedure can no longer provide a proof of contamination (due to non-independence and heuristic nature of the fitering step), their magnitude serves as heuristic evidence for contamination. The resulting p-values and empirical CDFs (Figure 4) of the 43 test sets are indicate mild deviation from the null hypothesis, consistent with the findings of mild test set contamination in Touvron et al. (2023).
Related Work
Our work relates to a large literature on data memorization, privacy, and membership inference attacks for large langauge models. We discuss some of the most relevant works to ours below.
There is a substantial literature studying memorization of data in large language models, often from the privacy perspective (Carlini et al., 2021; 2019; Kandpal et al., 2022; Mattern et al., 2023; Carlini et al., 2023). Most of these works have focused on analyses of what is memorized and whether private information can be extracted from a large langauge model, but do not build tests to specifically identify test set contamination. Our work has a narrower focus on test set contamination, but this also allows us to build tests that provide more precise guarantees of contamination.
Data contamination has been studied in many contexts, including in the study of pretraining corpora ((Dodge et al., 2021)) as well as in the analysis section of many language model papers (Hoffmann et al., 2022; Brown et al., 2020a; Gao et al., 2020). The n-gram based analyses in these papers can shed light on contamination, but they can have high false positives (e.g. SQuAD (Rajpurkar et al., 2016) containing Wikipedia) and are limited to those datasets chosen for analysis. Our approach enables third party tests of dataset contamination with only access to log probabilities, enabling broader testing, without having to trust the model provider.
For third-party tests of contamination, there have been a few recently proposed heuristics. Sainz et al. (2023) propose to identify test set contamination in GPT-3 and GPT-4 by prompting the models to generate verbatim examples from a test set. The contemporaneous work of Golchin & Surdeanu (2023) similarly proposes to identify contamination in black-box models by prompting a model to generate completions of random example prefixes and using GPT-4 to judge the closeness between the completion and the ground truth. While these approaches are efficient and do not require access to pre-training data, they do not enjoy the same provable false-positive guarantees of our work and require strong memorization that is detectable in generated outputs.
Closest to our work is the exposure statistic in Carlini et al. (2019) and other subsequent variations (Mattern et al. (2023)), which tests the perplexity differences between a target sequence and random sequences. The idea of comparing the rank of the target log probability to some baseline distribution is similar to our work. However, our work is distinct in using the exchangability of datasets to obtain an exact null distribution (giving us provable guarantees when identifying contamination) and in developing a sensitive and efficient shard-based test.
Beyond language modeling, identifying the presence of a particular set of examples in the training data of a machine learning model is related to the security and privacy topic of membership inference (Shokri et al. (2017); Mattern et al. (2023)). Our work contributes to this literature by developing a new form of membership inference attack that leverages the exchangability of benchmark datasets.
Limitations
We highlight a few limitations of our approach for detecting test set contamination. First, the p-values presented in this paper do not have multiple test corrections applied, as it is difficult to define the ‘total number of hypotheses’ tested throughout development.
Second, any application of this test in practice will likely involve taking an off-the-shelf benchmark dataset , for which it will be difficult to know if the dataset is truly exchangable. Heuristic negative controls such as our BioMedLM experiments can be helpful, but we cannot ever prove that a dataset is exchangable without knowing its data generating process. We strongly encourage future dataset creators to apply a random shuffle to their datasets (and to publicize this fact), which would allow our tests to be applied.
Finally, our tests focus on the case of verbatim contamination where a language model ingests a test set directly. Contamination can happen in many other ways, such as when a language model consumes a data source used in the construction of a benchmark (e.g. Wikipedia used in SQuAD, professional tests in MMLU). Verbatim memorization of a test set is not the only form of contamination, and our test cannot rule out the possibility of more complex forms of partial contamination.
Conclusion
In this work, we demonstrated that it is possible to construct a statistical test for test set contamination that provides false positive rate guarantees and requires nothing other than the ability to compute log probabilities. We construct new, sharding based tests for contamination and demonstrate their power on both carefully constructed canaries as well as publically available language models. We view these tests as a first step towards building powerful third party tests of contamination, and we believe it is an exciting open problem to build tests that are capable of reliably detecting contamination at the single-duplication-count regime.
Acknowledgements
We gratefully acknowledge support from the CRFM Levanter team, especially David Hall, for both computational resources and infrastructure support, and to Google’s TPU Research Cloud (TRC) for Cloud TPUs used in the pretraining experiments. Nicole Meister was supported by NSF GRFP DGE-2146755. Niladri Chatterji and Faisal Ladhak were supported by SAIL and Tatsunori Hashimoto was supported by a gift from IBM and a Hoffman-Yee grant. Finally, we would like to thank Nicholas Carlini and Percy Liang for insightful discussions on memorization and test set contamination.
References
Appendices
To compute log-likelihoods for sequences exceeding the context length, we use a strided window approach, with a stride equal to half of the model’s context length. We find that decreasing the stride beyond half the context length does not yield significant gains.
B Pretraining Details
We elaborate on the hyperparameters and training procedure of our 1.4B language model, trained from scratch on intentionally contaminated Wikitext.
We use a GPT-2 architecture with 1.4B parameters, with the architecture hyperparameters given by a hidden dimension of 1536, 24 heads, 48 layers, and a sequence length of 2048. The training batch size was 256. Based on the number of training tokens, sequence length, and training batch size, we trained this model for 46000 steps so as to consume the tokens in our mixture datasets exactly once. The model was optimized using AdamW with a learning rate of 1e-4 and weight decay of 0.1. We trained the model using Levanter on a v3-128 TPU instance on Google Cloud for 1.5 days (Hall et al. (2023)).
C 10 Canary Datasets
In this section we provide additional details on the 10 canary datasets we injected into Wikitext to form our pretraining data. For BoolQhttps://github.com/google-research-datasets/boolean-questions (Clark et al., 2019), HellaSwaghttps://rowanzellers.com/hellaswag/ (Zellers et al., 2019), MNLIhttps://cims.nyu.edu/~sbowman/multinli/ (Williams et al., 2018), Natural Questionshttps://github.com/google-research-datasets/natural-questions (Kwiatkowski et al., 2019a), TruthfulQAhttps://github.com/sylinrl/TruthfulQA/blob/main/data/finetune_truth.jsonl (Lin et al., 2022), PIQAhttps://yonatanbisk.com/piqa/ (Bisk et al., 2019), we sample a random subset of 1000 examples. For OpenbookQAhttps://allenai.org/data/open-book-qa (Mihaylov et al., 2018b), because of its smaller test set of size n=500, we used all 500 examples. Finally, for MMLUhttps://github.com/hendrycks/test (Hendrycks et al., 2021), we chose from subsets with no multi-line examples and having at least 500 examples, specifically Professional Psychology (n=611), MMLU Professional Law (n=1000), MMLU High School Psychology (n=544). Finally, we shuffle the examples in all datasets to make them exchangeable. In Table 3, we provide additional information about the injected datasets including number of examples, average words per example, and number of tokens per dataset. For each duplication rate, we included a short, medium and longer dataset, as measured by the total token count. The total token count of injected benchmarks is 19.235M tokens, meaning that the injected dataset is less than 0.1% of the entire pre-training dataset.
D Full MMLU Results
We list in table 4 the full set of MMLU results used to generate the omnibus p-values for contamination by MMLU listed in table 2, before filtering out for suspected non-exchangeability.