Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models

Yihong Dong, Xue Jiang, Huanyu Liu, Zhi Jin, Bin Gu, Mengfei Yang, Ge Li

Introduction

In recent years, LLMs have revolutionized the fields of natural language processing (NLP), artificial intelligence, and software engineering. To evaluate LLMs’ capabilities in various downstream tasks, such as automatic question answering, natural language reasoning, and code generation, people conduct extensive tests for LLMs based on enormous benchmark datasets Chen et al. (2021); Cobbe et al. (2021). The results indicate that LLMs exhibit superior performance on these tasks. While marveling at the powerful capabilities of LLMs, people usually want to determine whether an LLM’s excellent performance is due to the genuine understanding of tasks to achieve generalization, or merely because it has seen the test data to form memorization, i.e., suffering from data contamination.

Data contamination, also known as data leakage, refers to the scenario where the test data has been included in the model’s training data (Magar et al., 2022; Dickson, 2023), leading to the model performing exceptionally well on these leaked test data. Owing to the vast size and wide-ranging sources of the pre-trained datasets for LLMs, they are more susceptible to data contamination, which can be primarily categorized into two situations: 1) For existing benchmark datasets, they are more easily leaked because of massive text quotes, code reuse, and synthetic data in LLMs’ training data. 2) For upcoming benchmark datasets, newly constructed test data may already exist in the continuously evolving training data of LLMs since people are usually unaware of the specifics of LLMs’ training data. Consequently, it becomes formidable to prevent data contamination for LLMs.

Data contamination exerts a profound and deleterious impact on LLMs Zhou et al. (2023); Jacovi et al. (2023); Roberts et al. (2023). As shown in Figure 1, with LLMs continuing to learn on contaminated data (i.e., both leaked data and other training data), their performance keeps improving on leaked data but stagnates and even degrades on similar data. This example reflect that data contamination can lead to a substantial overestimation of models’ performance, thus affecting the trustworthiness and effectiveness of LLMs in practical applications. Furthermore, data contamination may also conceal the potential flaws of models, presenting major obstacles for people to identify and improve upon LLMs’ shortcomings. Therefore, it is crucial for LLMs to detect data contamination and ensure trustworthy evaluation.

Although acknowledged the significance, data contamination detection and trustworthy evaluation for LLMs still persist as open and challenging issues Yang et al. (2023); Huang et al. (2023). The difficulties of data contamination detection can be essentially attributed to three factors: 1) Opaque Training Data. It is usually non-public and comprehensive, while continuously evolving for new LLMs. 2) Black Box Models. The parameters and output probabilities of LLMs may not be available, such as ChatGPT and GPT-4 OpenAI (2023). 3) Proliferation of Synthetic Data. It could introduce the variantsThese variants may include, but are not limited to, translations into other languages, additions of explanations or intermediate processes, and provisions of alternate solutions. of test data to training data. Further, the evaluation to mitigate the impact of data contamination has hardly been studied.

In this paper, we overcome the preceding challenges by proposing CDD: Contamination Detection via output Distribution for LLMs. CDD uses the sampled texts to identify the peakedness of LLM’s output distribution for data contamination detection. We follow a hypothesis that training is likely to alter the model’s output distribution, resulting in a more peaked output distribution for training data, thereby tending the model towards specific outputs on these data. On this basis, we also present TED: Trustworthy Evaluation via output Distribution, which is designed to mitigate the impact of data contamination in evaluation by correcting LLM’s output distribution.

We construct two new datasets, i.e., DetCon and ComiEval, for data contamination detection and contamination mitigation evaluation tasks, respectively. Experimental results demonstrate that CDD achieves state-of-the-art (SOTA) performance and is also suitable for identifying variant contamination, i.e., existing the variants of test data in training data. TED successfully mitigates the impact of data contamination in evaluation across various scenarios. Furthermore, we also provide strong evidence that ChatGPT suffers from data contamination on HumanEval dataset.

Motivation Example

A powerful LLM that transcends memorization has the capability to generate diverse outputs in response to a given input. Considering the huge vocabulary size of LLMs, which encompasses a good number of tokens with analogous semantics, the output distribution sampled from LLMs ought to not exhibit peakedness. However, when LLMs solely form memorization via training, LLMs are prone to generate outputs that abnormally resemble their training data, as shown in Figure 2. From a statistical perspective, assuming that the average probability of LLM’s output tokens is 0.95, the likelihood of sampling two outputs that contain the same 100 consecutive tokens is about 0.005 < 0.01, which is an extremely improbable event. Therefore, if an LLM consistently outputs some identical or highly similar texts through sampling, it is most likely caused by memorization.

Figure 3 displays an example of how the LLM’s output distribution changes as the degree of data contamination varies. We model the LLM’s output distribution by computing the edit distances of sampled texts, referred to as edit distance distribution (§ \refEDD\S~{}\ref{EDD}). As shown in Figure 3, in the absence of data contamination (i.e., occurrence 0), the density of zero edit distance stands at 0.0035, where zero edit distance means that sampled texts exactly match. However, upon the LLM being exposed to the leaked data even once (i.e., occurrence 1) during training, the density of zero edit distance escalates sharply to more than 20 times larger than the original, showing the peakedness. Therefore, the impact of data contamination on the LLM’s output distribution is substantial.

In this paper, to the best of our knowledge, we are the first to consider from the standpoint of LLMs’ output distribution to address the challenges associated with data contamination detection and contamination mitigation evaluation, employing only the sampled texts without access to the output probability and training data.

Methodology

In this section, we first establish the edit distance distribution (§ \refEDD\S~{}\ref{EDD}), and then on this basis, we design CDD for data contamination detection (§ \ref3.2\S~{}\ref{3.2}) and TED for contamination mitigation evaluation (§ \ref3.3\S~{}\ref{3.3}). The illustration of our proposed approaches is presented in Figure 4.

Edit distance Levenshtein et al. (1966) is a measure of similarity between two strings, which is defined as the minimum number of operations required to transform one string into the other. The operations typically include insertion, deletion, or substitution of a single character.

Considering the generation of LLMs is based on tokens instead of characters, we adopt token-level edit distance in this paper. Given two strings aa and bb, token-level edit distance is calculated as:

where Len⁡(a)\operatorname{Len}(a) means the length of tokenized aa, Head⁡(a)\operatorname{Head}(a) means the first token of tokenized aa, Tail⁡(a)\operatorname{Tail}(a) means the string consists of all tokens of tokenized aa following Head⁡(a)\operatorname{Head}(a). We use dynamic programming to speed up calculations and rolling arrays to reduce space overhead.

Given an LLM, we can model its output distribution by computing the edit distances of sampled texts S={s1,s2,...,sn}S=\{s_{1},s_{2},...,s_{n}\}, where nn is the number of samples. Specifically, we define the density function ρ\rho as:

2 CDD for Data Contamination Detection

Given a test data {x,y}\{x,y\} consisting of a prompt xx and the corresponding answer yy, we aim to detect if this data has been trained by the model M\mathcal{M}.

We sample SS from M\mathcal{M} with the input xx to calculate ρ\rho. For data contamination detection, the calculation of ρ\rho can be simplified as:

However, ρ′(d)\rho^{\prime}(d) assumes that test data must be leaked in its original form {x,y}\{x,y\}, and does not take into account the possible contamination of the variant form, i.e., {x,y^}\{x,\hat{y}\}.

Through observation, we find that the copy percentage of model outputs increases as the degree of data contamination increases, as shown in Figure 1. Therefore, we approximate yy by the model’s output texts and finally choose to replace yy with the model’s greedy search text st=0s_{t=0}, which can be easily achieved by setting temperature tt = 0 when sampling. Thus,

In this work, we employ ρ∗(d)\rho^{*}(d) to measure edit distance distribution by default.

Further, we define the peakedness of edit distance distribution as

where FF is the cumulative distribution function, α∈\alpha\in is a hyper-parameter to control the similarity, and and ll is defined as:

Through identifying the peakedness, CDD can detect data contamination on test data as:

where ξ∈\xi\in is hyper-parameter to control the threshold. The pseudocode of CDD for data contamination detection is shown in Algorithm 1.

3 TED for Contamination Mitigation Evaluation

We achieve contamination mitigation evaluation using TED, which includes two rules to correct the LLM’s output distribution, i.e., exclude peakedness and remove duplicates.

1) Exclude Peakedness. We hope to restore the uncontaminated sampling results by excluding the peakedness in the LLM’s output distribution, while excluding the greedy text st=0s_{t=0} which is most likely to represent the leaked data potentially memorized by the LLM.

where τ∈[0,+∞)\tau\in[0,+\infty) is a hyper-parameter to control the difference.

2) Remove Duplicates. It aims to remove the duplicate sampling results, especially those differing from st=0s_{t=0}, which are also less likely to duplicately occur in the uncontaminated sampling results.

In the evaluation phase, an evaluation metric E\mathcal{E} using TED to mitigate the impact of data contamination can be defined as:

The pseudocode of TED for contamination mitigation evaluation is shown in Algorithm 2.

Experiment

In this section, we first introduce two datasets, DetCon and ComiEval, tailored for the tasks of data contamination detection and contamination mitigation evaluation, respectively (Section 4.1). We then evaluate the efficacy of CDD on the DetCon dataset (Section 4.2). Following this, we assess the performance of TED on the ComiEval dataset (Section 4.3). Finally, we demonstrate the application results of both CDD and TED in the real-world scenario (Section 4.4).

Considering the absence of datasets for data contamination detection and contamination mitigation evaluation tasks, we dedicate more than 2100 hours to constructing the DetCon and ComiEval datasets, utilizing two A6000 GPUs (48GB ×\times 2).

To simulate different data contamination scenes of LLMs, we conduct extensive experiments that span two representative downstream scenarios, two contamination forms, four mixing numbers, three learning rates, and twenty-one occurrences (i.e., contamination degrees) on four different base LLMs around 7B. The detailed statistics can be found in Table 1. We employ LoRA Hu et al. (2022) to fine-tune the base models on these various settings On this basis, we construct the DetCon and ComiEval datasets.

contains 2224 data contamination detection tasks, covering two scenarios (code generation and logical reasoning) and two contamination forms (original and variant), which need to detect whether a specific LLM has contamination on a particular data. We randomly select the data from the leaked dataset and the LLM from the settings in Table 1, where occurrence 0 refers to ‘uncontaminated’ and the others denote ‘contaminated’.

ComiEval

contains 560 contamination mitigation evaluation tasks, consisting of a randomly selected contaminated model from Table 1 and the corresponding uncontaminated model, which need to evaluate the performance of the contaminated model and try to mitigate the impact of data contamination to approach the performance of the uncontaminated model.

The detailed statistics and introductions of DetCon and ComiEval datasets can be found in Appendix A.

2 Data Contamination Detection

Experimental Setup. We compare CDD with baselines, including 1) N-gram: We employ widely-used 13-gram for both char-level and token level; 2) Embedding Similarity: Use the embedding of the base model to compute similarity; 3) Perplexity: Compute the perplexity of the original answer given the prompt; 4) Min-k% Prob: Compute the minimum k% probability of the original answer given the prompt, and 5) LLM Decontaminator: Use other LLM to determinate the similarity and we employ ChatGPT as this LLM. The differences between CDD and baselines are shown in Table 2. For hyper-parameters, we set α=0.05\alpha=0.05, ξ=0.01\xi=0.01, the cap of ll as 100 for CDD by default, and baselines follow the settings in their paper.

As presented in Table 3, compared with other contamination detection approaches, CDD attains SOTA performance in both code generation and logic reasoning scenarios. CDD exhibits steady improvements across the Accuracy, F1 Score, and AUC metrics, with the average relative improvement ranging between 21.8% and 30.2%. Moreover, the advantage of CDD is that it only requires the sampled texts of LLMs to detect data contamination, without the need for additional conditions in Table 2.

We also evaluate the performance of CDD and other contamination detection approaches in identifying original and variant forms of data contamination, as shown in Figure 5. In the cases of original contamination, as the degree of contamination increases, the detection effectiveness across all approaches improves. CDD outperforms the other approaches at lower contamination degrees, which are more challenging to detect. In contrast, in the cases of variant contamination, CDD alone maintains robust performance, whereas the other approaches encounter significant limitations.

We fix the hyper-parameter α\alpha and ξ\xi intuitively for CDD in the experiments. In Figure 4 (a) and (b), we analyze the influence of α\alpha and ξ\xi empirically on DetCon dataset by changing itself and fixing another hyper-parameter. The results indicate that there is still room for further improvements with the better hyper-parameter setup of α\alpha and ξ\xi.

3 Contamination Mitigation Evaluation

We evaluate the effectiveness of TED for contamination mitigation in different learning rates, base LLMs, mixing ratios, contamination forms, and occurrences on ComiEval. We set the hyper-parameter γ=2\gamma=2 and use Pass@1 Chen et al. (2021) as the evaluation metric E\mathcal{E}.

Effect of TED.

TED can steadily mitigate the performance improvements across different settings and occurrences in data contamination scenes, as shown in Figure 6. Moreover, the advantage of TED is that the performance influence of TED on the uncontaminated model (i.e. 0 occurrences) is small and almost negligible. However, as contamination degrees continue to increase, the performance influence of TED becomes apparent in all of the different settings.

We analyze the effects of each component in TED, as shown in Table 4. The main function is provided by the rule of exclude peakedness, followed by the rule of remove duplicates. Both components are beneficial to TED and are also effective when employed alone.

As illustrated in Figure 4 (c), an increase in the hyperparameter γ\gamma leads to a more pronounced suppression of performance improvements attributable to data contamination by TED. Meanwhile, it also marginally decreases the performance of the uncontaminated model.

4 Real-World Application

In real-world applications, we apply CDD and TED for ChatGPT and construct two new datasets to assist evidence: 1) CodeForces2305 comprises 90 of the easiest level programming problems collected from the CodeForces website since May 2023, which is after the most recent update deadline of ChatGPT’s training data, i.e., April 2023. 2) HumanEval_R is reconstructed on HumanEval, which replaces its function signature, translates its requirements into German, French, and Chinese, selects different public test cases from the work Dong et al. (2023a) to prompt, and remains the private test cases for testing. To enhance the detection precision, we set the hyper-parameters α\alpha to 0 and ξ\xi to a larger value of 0.2 for CDD. We keep γ\gamma at the default value of 2 for TED.

Data contamination for ChatGPT.

As shown in Table 5, on HumanEval dataset, both two versions of ChatGPT exhibit high Avg. Peak and Leak Ratio, and the ‘post’ version become higher as ChatGPT continues to learn on new data. Considering the implementation of more stringent α\alpha and ξ\xi, it is posited that ChatGPT is likely to suffer from data contamination on HumanEval dataset and become more serious over time. This hypothesis is further evidenced through evaluations conducted on HumanEval_R and CodeForces2305 datasets. HumanEval_R indicates their high Avg. Peak and Leak Ratio are not easily attributable to the difficulty of problems. By modifying prompt forms through a process of reconstruction, all of the performance, Avg. Peak, and Leak Ratio of ChatGPT are significantly reduced. On CodeForces2305 dataset, which is unlikely to be involved in data contamination, ChatGPT’s performance was markedly lower than anticipated, with the Avg. Peak at less than 0.01 and Leak Ratio of 0. Moreover, TED demonstrates significant effectiveness on both the contaminated HumanEval and HumanEval_R.

Related Work

The concept of data contamination for LLMs can be derived from the context of GPT-3 Brown et al. (2020). Due to the vastness of the pre-training corpus of GPT-3, it inevitably overlapped with some evaluation benchmarks. Therefore, GPT-3 adopted 13-gram overlap detection to remove the data in the training set that conflicts with the test set of benchmarks.

Some work Pan et al. (2020); Zhou et al. (2023); Jacovi et al. (2023); Dodge et al. (2021) exposed the serious consequences of data contamination and urged attention to this problem. However, most currently released LLMs did not open their pre-training corpus, which poses a new challenge for data contamination detection. Recent work tried to detect contamination without access to the pre-training corpus Oren et al. (2023); Deng et al. (2023); Golchin et al. (2023). Min-k% Prob Shi et al. (2023) calculated the average of the k% smallest probabilities of generated tokens and considered it as contaminated if it exceeded a certain threshold. The work Li (2023) assumed that data leaked into the training set tends to exhibit lower perplexity and utilizes perplexity analysis for detection. However, they often require other model outputs (e.g. probability) in addition to text, presenting challenges in detecting closed-source LLMs like ChatGPT, and they ignore the potential contamination from variants of test data.

Recent investigations Huang et al. (2023); Yang et al. (2023) have suggested that filtering training data based on n-grams may not effectively address the issue of data contamination, especially concerning semantically equivalent sentence rephrasing. To this end, LLM Decontaminator Yang et al. (2023) detected the similarity of test data and training data based on other advanced LLMs.

Our work requires only sampled texts to detect LLM’s data contamination via output distribution and considers the potential variant contamination.

Contamination Mitigation Evaluation.

To mitigate the impact of data contamination and ensure trustworthy evaluations, several approaches focus on constructing new evaluation benchmarks Golchin et al. (2023). The work Zhu et al. (2023) employs an LLM to paraphrase the contaminated dataset for evaluations. However, LLM’s synthetic data is widely used for training, which already contains lots of paraphrased data Yang et al. (2023). The work Li et al. (2023b) leverages temporal information to construct a benchmark beginning from January 2023. However, building a high-quality benchmark is costly and time-consuming, and unfortunately, the training data deadline for ChatGPT and GPT-4 has been updated from September 2021 to April 2023 and continues to be delayed.

Our work achieves contamination mitigation evaluation from the standard of LLM’s output distribution and is orthogonal to the preceding works.

Conclusion

In this paper, we have proposed two novel approaches, namely CDD and TED, to deal with data contamination detection and contamination mitigation evaluation for LLMs, considering the LLM’s output distribution. We construct two corresponding datasets, i.e., DetCon and ComiEval, for these two tasks. Extensive experimental results indicate the superiority and versatility of CDD and TED. Moreover, we also discover that ChatGPT is likely to suffer from data contamination on HumanEval dataset. We hope to shed light on this direction and call more attention to data contamination issues.

Limitations

Our work has several limitations, which we aim to address in our future work:

First, the validation of our work is mainly focused on benchmarks for code generation and logical reasoning, which are highly representative and widely adopted. In the future, we will further validate our approaches on other benchmarks.

Second, our approaches require multiple samplings to compute the output distribution, and the more samplings conducted, the better the effect. We can use parallel sampling techniques to speed up sampling, thereby reducing time overhead.

Third, considering the limitation of computational resources, we employ a popular parameter-efficient fine-tuning approach, i.e., LoRA, instead of full-parameter fine-tuning to simulate data contamination for LLMs. In future work, we plan to attempt full-parameter fine-tuning.

Finally, in constructing our datasets, we assume that the base LLMs used do not suffer from data leakage on the selected benchmarks. However, in reality, these LLMs may have slight data contamination. To completely avoid this issue, it might be necessary to retrain an LLM from scratch on a training set known to be entirely free of test data. However, undertaking such a process would be prohibitively costly.

References

Appendix A Details of Dataset Construction

Test Data. We choose the HumanEval Chen et al. (2021) dataset for code generation and the GSM8K Cobbe et al. (2021) dataset for logical reasoning.

LLMs. For code generation tasks, we use CodeLlama-7B Rozière et al. (2023) and CodeGen-6.7B Nijkamp et al. (2023); for natural language processing tasks, we select Llama2-7B and Bloom-7B.

Training Data. Code generation tasks use the training data from StarCoder Li et al. (2023a), while natural language processing tasks use RedPajama Computer (2023).

Next, we construct the dataset DetCon for data contamination detection, starting with the construction of uncontaminated samples. We directly use the outputs generated by LLMs on the test data, representing uncontaminated data. Then, we construct contaminated samples by simulating different contamination scenes:

Explicit and Implicit Contamination. Explicit contamination refers to the direct use of test data for training, while implicit contamination refers to training with variants of the test data.

Proportion in the training data. We use different amounts of training data mixed with test data to train LLMs. The proportions of the test dataset mixed with training data include 1:0, 1:0.1k, 1:1k, 1:10k.

Different learning rates. Considering the effect of learning rate on model training, we chose three different learning rates: 1e-3, 2e-4, and 4e-8.

Degree of data contamination. Training LLMs with contaminated data for more epochs indicates a higher degree of contamination. Epochs range from 0 to 20, where 0 means no training of LLM, indicating no contamination, reserved for constructing uncontaminated samples.

By combining these four different scenes, we can construct a variety of composite data contamination scenes. For each piece of test data, we randomly select the results generated by LLMs under one of the contamination scenes as the contaminated samples. Following the previous works Jiang et al. (2023); Dong et al. (2023b, c); Li et al. (2024), in generating these samples, we also record the outputs of greedy search with a temperature parameter of 0 (1 sample) and 50 samples obtained by sampling with a temperature of 0.8.

This construction approach aims to comprehensively cover possible data contamination scenes, ensuring we can accurately assess the performance of LLMs in the face of different types and degrees of data contamination.

Finally, we construct the dataset ComiEval for contamination mitigation evaluation. We selected 560 LLMs and their generated outputs from all the constructed contaminated LLMs as the task inputs. Then, we used the performance of the corresponding uncontaminated LLMs on these two test datasets as the target output, serving as the evaluation criterion. Table 6 demonstrates the statistics of DetCon and ComiEval, respectively.