Don't Make Your LLM an Evaluation Benchmark Cheater
Kun Zhou, Yutao Zhu, Zhipeng Chen, Wentong Chen, Wayne Xin Zhao, Xu Chen, Yankai Lin, Ji-Rong Wen, Jiawei Han
Introduction
Goodhart’s Law: “When a measure becomes a target, it ceases to be a good measure.”
Large language models (LLMs) have achieved remarkable success across a variety of real-world applications Brown et al. (2020); Zhao et al. (2023); Zhu et al. (2023). By pre-training large Transformer models on massive text corpora, LLMs can possess excellent task-solving capacities, i.e., using zero-shot or few-shot prompting Brown et al. (2020). To better understand how LLMs evolve in model capacity, it becomes essential to construct reliable evaluation benchmarks to test the ability level of LLMs in various tasks, e.g., knowledge reasoning and math problem solving.
Recently, a surge of high-quality evaluation benchmarks Hendrycks et al. (2021); Huang et al. (2023) have been proposed to provide a comprehensive capability evaluation of LLMs. Typical benchmarks include MMLU Hendrycks et al. (2021) (for measuring multitask language understanding ability), Big-Bench Srivastava et al. (2022) (for quantifying and extrapolating the capabilities of LLMs), and AGIEval Zhong et al. (2023) (for evaluating the abilities of tackling human-level tasks). These benchmarks have made great efforts in creating or collecting test resources for evaluating the performance of LLMs. Based on these benchmarks, one can conveniently examine the effect of new training strategies or monitor the training status of LLMs (either pre-training or supervised fine-tuning). It has become common to report the results on these evaluation benchmarks for demonstrating the effectiveness of newly released LLMs OpenAI (2023); Touvron et al. (2023b); Anil et al. (2023). Furthermore, to compare the performance of different LLMs, various leaderboards have been also created to rank LLMs according to their performance on existing or new evaluation benchmarks, such as OpenCompass Contributors (2023) and C-Eval Huang et al. (2023).
Despite the wide use of these benchmarks and leaderboards, increasing concerns Aiyappa et al. (2023); Li (2023) are growing about the fairness and reliability in evaluating existing LLMs. A major issue is that the data contamination or leakage is likely to occur for large-scale benchmark evaluation, which means that LLMs are trained with relevant or exactly the same data for test. Such an issue could be unconsciously triggered, since we might be unaware of the future evaluation datasets when preparing the pre-training corpus. For example, GPT-3 has found that Children’s Book Test dataset Hill et al. (2016) was included in the pre-training corpus, and LLaMA-2 has mentioned that the contexts in BoolQ dataset Clark et al. (2019) are extracted verbatim from the webpages, which may be included in the publicly available corpus.
Indeed, when conducting evaluation with existing benchmarks, the results of evaluated LLMs are mostly obtained by running them on local servers or via API calls. During this process, there is no strict checking on any potentially inappropriate ways (e.g., data contamination) that would cause an unnormal improvement of evaluation performance. To make matters worse, the detailed composition (e.g., data sources) of the training corpus is often regarded as the core “secret” of existing LLMs. Therefore, it becomes difficult to directly examine the contamination issues when performing the evaluation for benchmark maintainers.
Considering this issue, the aim of this paper is to draw attention on appropriately using existing evaluation benchmarks and avoiding any misleading behaviors in obtaining or interpreting the evaluation results. Specifically, we mainly focus on discussing the potential effect of benchmark leakage, which refers to the case that test data or relevant data (e.g., training set) has been included in the pre-training corpus. It would cause an unfair performance advantage when comparing different LLMs or assessing the ability level of some specific LLMs. As we discussed before, this issue tends to become increasingly more common as we try to collect more public text data for training. To investigate this issue, we set up several benchmark leakage settings that should be totally avoided during evaluation, including the leakage of training sets, test prompts, and test sets. Based on the three settings, we continually train four popular language models, ranging from 1.3B to 7B, and test the performance of the four models on a number of existing benchmarks. In addition, we also examine the potential risk of benchmark leakage on other abilities.
The experimental results reveal that benchmark leakage can lead to an unfair boost in the evaluation performance of LLMs. Smaller LLMs (e.g., a 1.3B model) can be deliberately elevated to outperform larger models on certain tasks. As a side effect, the performance of these specially trained LLMs on other normally tested tasks would likely be adversely affected if we fine-tune or train the model only with these leaked data.
By examining the potential risks of benchmark leakage, we would like to emphasize the importance of fair and appropriate evaluation for LLMs, and propose several suggestions to improve the evaluation for LLMs:
As general suggestions, more benchmarks from diverse sources, covering both basic ability (e.g., text generation) and advanced ability tests (e.g., complex reasoning), should be used for comprehensively estimating the capabilities of LLMs.
As suggestions for LLM developers, it is important to perform the data decontamination checking between pre-training data and any related data (e.g., training and test sets) when using evaluation benchmarks. In addition, it is also necessary to report the contamination analysis on the evaluated benchmarks as reference. We also suggest reporting the detailed composition of the pre-training data.
As suggestions for benchmark maintainers, we suggest that a diverse set of test prompts should be employed for reducing the influence of the prompt sensitivity. It is also meaningful to conduct the contamination analysis between the benchmark data and existing pre-training corpus, alerting any potential contamination risks. For evaluation, each submission is suggested to be accompanied with a special contamination analysis report.
Empirical Study about Benchmark Leakage
During pre-training, the data contamination or leakage about possible evaluation benchmarks, is likely to be unconsciously triggered Oren et al. (2023); Sainz et al. (2023). It would violate regular evaluation settings for assessing zero/few-shot generalization capability, thus affecting the capability assessment of LLMs. To better understand the potential influence of the benchmark leakage issue, we conduct an empirical study that continually trains small-sized LLMs on three settings with different levels of information leakage.
Our empirical study aims to test the influence of possible benchmark leakage issues on the evaluation results of LLMs. A benchmark typically contains a set of test examples, and relies on fixed templates to prompt LLMs for evaluation. Such an evaluation process may lead to three types of benchmark leakage risks, that is, including (1) test prompt, (2) test set, or (3) other relevant data (e.g., training set) into the pre-training corpus. Considering the above settings, we simulate three extreme leakage issues where the three types of information have been used for continually training LLMs, and design the following evaluation settings.
Using MMLU Training Set: the auxiliary training set provided by the official MMLU benchmark Hendrycks et al. (2021) is used for training.https://github.com/hendrycks/test. The auxiliary training set contains data collected from several question-answering benchmarks such as ARC, OBQA, and RACE.
Using All Training Sets: in addition to MMLU training set, the training sets of all other collected evaluation benchmarks are also used for training (details are provided later).
Using All Training Sets with Test Prompt: all the training sets, with their corresponding test prompts, e.g., task description and few-shot demonstration, are used for training.
Using All Training and Test Sets with Test Prompt: all the training sets, test prompts, and test sets of all the collected evaluation benchmarks are used for training. (CAUTION: this is the most extreme case, where all information is leaked. We conduct this experiment only for reference, and this should never occur.)
Evaluation Benchmark
To make the empirical study, we select the widely-used benchmark MMLU and employ a number of question-answering (QA), reasoning, and reading comprehension datasets for evaluation.
MMLU: it has become one of the most commonly used evaluation benchmarks for LLMs’ ability of world knowledge possessing and problem solving. It covers 57 tasks requiring diverse knowledge, such as math, history, science, and law. We report the 5-shot evaluation performance.
Open-domain QA Tasks: we select seven open-domain QA datasets where LLMs should answer the question solely based on intrinsic knowledge. We report the accuracy of LLMs under the zero-shot setting, i.e., BoolQ Clark et al. (2019), PIQA Bisk et al. (2020), Hellaswag Zellers et al. (2019), WinoGrande Sakaguchi et al. (2020), ARC Easy and Challenge Clark et al. (2018), OpenBookQA Mihaylov et al. (2018).
Reasoning Tasks: we select a commonsense reasoning dataset CommonsenseQA Talmor et al. (2019), and two commonly-used mathematical reasoning datasets GSM8k Cobbe et al. (2021) and AQuA Ling et al. (2017) for evaluation. We use chain-of-thought prompting and reuse the prompts provided by Wei et al. (2022) for evaluation and report the accuracy of LLMs.
Reading Comprehension Tasks: we select three English datasets RACE-Middle and RACE-High Lai et al. (2017), CoQA Reddy et al. (2019) and two Chinese datasets CMRC2018 Cui et al. (2019) and C3-Dialog Sun et al. (2020). As reading comprehension datasets have one paragraph and several QA pairs in a sample, we only test the accuracy of the last question and regard the paragraph and other QA pairs as the prompt. We report accuracy under the zero-shot setting for C3-Dialog, and utilize similar evaluation settings as GPT-3 Brown et al. (2020) for other tasks.
Backbone LLMs
To thoroughly analyze the effect of benchmark leakage on the evaluation performance, we select the following models for evaluation, which have provided pre-training details or conducted careful data contamination analysis.
GPT-Neo-1.3B Black et al. (2021): it is a Transformer-based model with GPT-3 architecture, pre-trained on the Pile Gao et al. (2021) dataset.
phi-1.5 Li et al. (2023): it is a 1.3B model trained on “textbook quality” data of 27B tokens, and can achieve comparable performance as much larger models.
OpenLLaMA-3B Geng and Liu (2023): it is an open-source project to reproduce LLaMA model with a permissive license, pre-trained on RedPajama dataset Computer (2023) of over 1.2T tokens.
LLaMA-2-7B Touvron et al. (2023b): it is an updated version of LLaMA Touvron et al. (2023a). It has been pre-trained on a mixture of publicly available online data of 2T tokens.
2 Results and Analysis
We report the evaluation results of LLMs after training with the benchmark leakage settings in Table 1 and Table 2. Overall, different levels of data leakage result in inflated model performance on benchmarks. We have the following observations.
First, we can see that using MMLU training set can greatly boost the evaluation results on the MMLU benchmark. However, this improvement comes at the cost of decreased performance on tasks unrelated to MMLU, (such as HellaSwag and GSM8k about commonsense and mathematical knowledge, respectively), suggesting that over-emphasizing a specific task may lower the model generalization capability. Besides, when incorporating all the training sets of the evaluated benchmarks, there is a notable performance increase across almost all the evaluated tasks. Incorporating training data converts the original zero/few-shot evaluation into an in-domain test task, making it easier for LLMs to achieve higher results. An intriguing finding occurs when we examine the result on the Chinese benchmark C3-Dialog. Despite the pre-training corpus of the four LLMs containing very little Chinese data, using training sets doubles their evaluation scores, e.g., elevating GPT-Neo-1.3B’s score from 24.18 to 48.62. This observation underscores the significance of avoiding training set leakage in pre-training, as it can lead to spurious performance improvements that distort the real assessment of model capabilities.
Second, the evaluation scores continue to rise as the data leakage becomes more severe. Remarkably, when the test prompts were leaked, smaller LLMs can even surpass much larger LLMs that were not trained with leaked data, e.g., “phi-1.5-1.3B+All Train S+Test P” outperforms LLaMA-65B on RACE-M (55.80 vs. 53.00) and RACE-H (52.82 vs. 48.00). This highlights the significance of the test prompt as valuable information from the evaluation benchmark, since it contains the detailed input format during test. During training LLMs, it is suggested to avoid such special learning with test prompts. Furthermore, this observation raises concerns about the robustness of using fixed test prompts in the evaluation benchmark, as it may not be resilient to the aforementioned leakage risk.
Finally, for reference, we examine the most extreme case where all test sets are leaked. The results are highlighted in grey font. As can be seen from these results, test data leakage significantly inflates benchmark performance, leading 1.3B LLMs to outperform 65B LLMs across most tasks. Evidently, this increase does not imply any improvement in capacity, but rather benchmark cheating.
Overall, benchmark leverage directly leads to an unfair advantage in evaluation results of the involved models, which should be strictly avoided when conducting any evaluation.
Potential Risk of Benchmark Leakage
In addition to the inflated performance that undermines the reliability of capability estimation, we also investigate whether the benchmark leakage issue would lead to potential risks in model capacity. Limited by the training compute, we can not conduct an exact checking that directly includes leakage data in pre-training data. Instead, we continually pre-train the LLMs on the training sets of all the selected evaluation benchmarks as in Section 2, without the mixture of any other data. Such a way is the most direct way for benchmark cheating (should be avoided). We speculate that it is likely to affect the capacities of LLMs on normally tested tasks (those without data leakage), due to “catastrophe forgetting” Luo et al. (2023); Goodfellow et al. (2013).As it is a very extreme scenario for simulation, we only employ it to explore the possibility of the subsequent impact when benchmark leakage occurs. The experiment procedure should be totally avoided in real training and evaluation.
After training on the leaked benchmark data, it would potentially mislead LLMs to overemphasize the specific knowledge and output style of the benchmark data, thereby potentially affecting their performance on other tasks. In this part, we conduct empirical experiments to examine the side effect on the model performance of other tasks.
To validate the effect, we select three tasks that are not involved in the leaked training data, consisting of two text generation tasks, i.e., LAMBADA Paperno et al. (2016) and XSum Narayan et al. (2018), and a code synthesis task HumanEval Chen et al. (2021) to evaluate LLMs in the zero-shot setting. LAMBADA is a language modeling task that tests the ability of LLMs to predict the last word based on the context, and we report the accuracy in predicting words. XSum, on the other hand, is a text summarization task that requires LLM to summarize the key information from long documents. For this task, we report the ROUGE-L metric, which measures the quality of the generated summaries by comparing them with the ground-truth summaries. For HumanEval, we adopt pass@10 as the evaluation metric.
Results Analysis
We show the results of LLMs with and without benchmark leakage on the three evaluation tasks in Table 3. First, we can observe that after training on the leaked data, the performance of all LLMs degrades on the two text generation datasets. Specifically, for OpenLLaMA-3B and LLaMA-2-7B, their text summarization abilities seem to be weakened after training on the leaked data, resulting in Rouge-L scores of 0.19 and 0.25 in XSum, respectively. Besides, by comparing the performance on HumanEval, we also see that data leakage primarily leads to performance degradation of LLMs in the code synthesis task.
This demonstrates that benchmark leakage may have a negative impact on the performance of these normally tested tasks (without data leverage).
2 Effect on Model Adaptation
After training on the leaked data, LLMs are trained to be specially fit for the benchmark data. However, LLMs might need to be further fine-tuned for attaining some specific goals (e.g., solving new tasks or serving emergent applications). In this part, we examine how inappropriately trained LLMs perform for subsequent adaptation.
To investigate the influence of data leakage on LLMs’ adaptation capability, we select two representative instruction datasets, i.e., Alpaca Taori et al. (2023) and CodeAlpaca Chaudhary (2023). Both of these datasets are synthetic and generated using the Self-Instruct method. For comparison, Alpaca primarily contains natural language instructions, whereas CodeAlpaca focuses on code generation instructions. We use these datasets to fine-tune the LLMs with or without training on the leaked data, and subsequently evaluate their performance on the previously mentioned text generation and code synthesis tasks.
Results Analysis
In Table 4, by comparing the performance of the instruction-tuned LLMs (+Alpaca or +CodeAlpaca) with and without training on the leaked data, we can see that the models with benchmark leakage still underperform their non-leaked counterparts. For the HumanEval dataset, the performance improvements of instruction tuning for LLMs trained with leaked data only reach approximately 80% of those achieved by models that are not trained on leaked data.
This indicates that benchmark leakage may lead to a decline in adaptation capability, constraining the LLMs’ ability to adapt or improve through subsequent fine-tuning processes. Note that this finding is derived when we fine-tune LLMs only with the leaked data. To enhance the current findings, it is also meaningful to conduct experiments that either include leaked data into pre-training data or mix leaked data with other instruction data. However, since our main purpose is to reveal that benchmark leverage might cause severe side effects on LLMs in addition to spurious performance improvement, we omit these experiments due to the compute limit.
Discussion
In light of the potential risks of benchmark leakage, it is necessary to revisit the existing evaluation settings for LLMs and investigate possible strategies to avoid such data contamination issues.
Based on our empirical findings in previous sections, the evaluation results of LLMs in specific benchmarks can be dramatically boosted when the related or same data of the test tasks is accidentally used for training. In the literature of machine learning, zero/few-shot learning often refers that the samples at test time were not observed during training for a learner Wang et al. (2021); Xian et al. (2019). It is evident that benchmark leverage does not comply with this requirement, making it unfair to compare different LLMs when such a case exists. Furthermore, data leverage can also bring an unfair advantage in the few-shot setting since the learner can observe more task-relevant data at training time.
In case of data leakage, the original zero-shot/few-shot generalization task would degenerate into much easier in-domain evaluation tasks, and it would intensify the phenomenon of benchmark hacking, i.e., a benchmark is no longer useful for evaluation due to the high performance of the involved comparison methods.
However, in practice, it is challenging to fully eliminate the leakage risk from model training Golchin and Surdeanu (2023); Shi et al. (2023). It is because an evaluation benchmark is often conducted based on some public text sources, e.g., webpages and scientific papers. In this case, the related data (e.g., the original text used to generate the test problems) might be occasionally included in the pre-training data of LLMs. Although existing evaluation datasets are easy to be excluded from pre-training data for training new LLMs, it is still difficult to identify all potential data dependencies between evaluation benchmarks and pre-training corpus. Such a test set contamination problem has been already noted in black-box language models Oren et al. (2023).
2 Suggestion for LLM Evaluation
Based on these discussions, we propose the following suggestions to improve existing capacity evaluation for LLMs.
Considering the potential risk associated with benchmark leakage, we recommend the use of a broader range of benchmarks from diverse sources for performance evaluation. This can help mitigate the risk of inflated results due to data contamination. If feasible, incorporating manual evaluation and conducting qualitative analysis would be also beneficial.
In addition to evaluating the advanced capabilities of LLMs (such as reasoning and factual knowledge), it is also necessary to perform evaluations on other datasets that focus on basic abilities, such as text generation. This comprehensive approach is necessary for a thorough estimation of LLMs’ capabilities.
Suggestions for LLM developers:
Perform strict checking on data decontamination in pre-training data to avoid any subsequent evaluation data being included during training. To achieve this, the -gram (generally, ) hash algorithm can be applied to examine the overlap between pre-training data and evaluation data of some specific task.
If possible, we suggest also excluding training data of mainstream evaluation benchmarks from pre-training data.
Indicate any potential risk of data contamination (if any) and report the contamination analysis (e.g., overlap statistics) when you present the results on some evaluation benchmark. An example can be seen in Llama-2’s report Touvron et al. (2023b).
Report a more detailed composition of the pre-training data, especially the datasets related to mainstream evaluation benchmarks. It is an important reference for checking the potential data leakage risk by the public audience.
Suggestions for benchmark maintainers:
Provide the detail of the data source for constructing the benchmark, and conduct the contamination analysis of the current dataset with mainstream pre-training corpora (as many as possible). The benchmark should explicitly alert possible contamination risks for commonly used pre-training datasets.
Each submission is suggested to be accompanied with a specific contamination analysis report from the result provider, where it can perform semantic relevance checking (e.g., overlap statistics) between pre-training data and evaluation data (both training and test data).
Provide a diverse set of prompts for testing. The final evaluation results should be averaged over these multiple runs. It can help reduce the sensitivity of specific prompts, and enhance the reliability of the model results.
Conclusion
In this paper, we conducted empirical studies to investigate the penitential risk and impact of benchmark leakage on LLM evaluation. We found that data leakage can largely boost the benchmark results of LLMs (even small models), making the evaluation unfair and untrustworthy. These findings suggest that such attempts should be strictly avoided for fairly assessing the model performance on evaluation benchmarks.
Despite that this issue is hard to be fully eliminated from the pre-training stage, we suggest several useful guidelines to improve the use of existing evaluation benchmarks. A key point is that both LLM developers and benchmark maintainers should be aware of the data contamination issue when interpreting and using the results from the performance leaderboards. In practice, several heuristic strategies can be useful to detect such potential contamination issues, e.g., calculating the token overlap between training and evaluation data. Besides, we also suggest benchmark test should be conducted with multiple task prompts for deriving a more stable and reliable model performance.
This work aims to draw the attention of the research community to the appropriate use of existing evaluation benchmarks for LLMs. More meaningful work can be conducted following this line, e.g., alerting the potential contamination datasets.
Limitation
In this work, we conducted preliminary experiments to emphasize the potential risks associated with benchmark leakage in training LLMs. However, there are still several limitations in our study.
First, our experiments involved continually training existing pre-trained LLMs with leaked data. We do not have sufficient computational resources to investigate the impact when directly incorporating benchmark leakage during the pre-training process. Given that the pre-training dataset is significantly larger than the benchmark data, introducing data leakage during pre-training might yield different findings. Nonetheless, we strongly recommend avoiding this situation as it would breaks the nature of zero-shot/few-shot evaluation.
Second, we did not explore more fine-grained data leakage scenarios in this study, such as only leaking training examples without labels and varying the proportion of the leaked dataset. We encourage more research efforts into this issue with more systematic studies.
Third, we did not calculate the degree of contamination between the mainstream benchmarks and commonly-used pre-training datasets, which could serve as an important reference for alerting LLM developers to adjust their evaluation settings. While we suggest that developers and benchmark maintainers report contamination analyses, accurately and efficiently estimating the contamination risk of each example in the benchmark is also a challenging task. For example, the suggested -gram hash algorithm may not detect semantic-level knowledge leakage risks.