LLM Evaluators Recognize and Favor Their Own Generations
Arjun Panickssery, Samuel R. Bowman, Shi Feng
Introduction
Self-evaluation is becoming a prominent part of the large language model (LLM) lifecycle. In methods like reward modeling (Leike et al., 2018; Stiennon et al., 2020), model-based benchmarks (Shashidhar et al., 2023; Zeng et al., 2023; Yuan et al., 2023; Fu et al., 2023; Li et al., 2024), self-refinement (Saunders et al., 2022; Madaan et al., 2023; Lee et al., 2023; Shridhar et al., 2023), and constitutional AI (Bai et al., 2022), LLMs are increasingly used to provide assessment, supervision, and oversight for themselves and other LLMs. LLM evaluators are shown to be highly accurate at approximating human annotators on various tasks, and are significantly more scalable (Hackl et al., 2023).
In self-evaluation, as the name suggests, the same underlying LLM acts as both the evaluator and the evaluatee. As a result, the neutrality of the evaluator is in question, and the evaluation can suffer from biases where the LLM evaluators diverge from humans in systematic ways (Zheng et al., 2024; Bai et al., 2024). One such bias is self-preference, where an LLM rates its own outputs higher than texts written by other LLMs or humans, while human annotators judge them as equal quality. Self-preference has been observed in GPT-4-based dialogue benchmarks (Bitton et al., 2023b; Koo et al., 2023), as well as for text summarization (Liu et al., 2023).
Towards understanding and mitigating self-preference, we study self-recognition—an LLM’s capability of recognizing its own outputs. We ask: Is self-preference truly self-preference, in the sense that the LLM prefers a text because it was generated by itself?
We measure their correlation while using prompting and fine-tuning to alter the LLM’s self-recognition capability. In order to provide signals for the causal link between self-recognition and self-preference, we also fine-tune the LLM on a comprehensive set of potential confounding properties.
Frontier LLMs exhibit self-preference in self-evaluation. On two summarization tasks, LLMs (GPT-3.5 Turbo, GPT-4, and Llama 2) disproportionately favor summaries written by themselves over those by other LLMs and from humans.
LLMs have non-trivial self-recognition capability out of the box. All three LLMs we evaluate achieve over accuracy at distinguishing their own outputs from other sources using simple prompts without fine-tuning. GPT-4 is accurate distinguishing itself from two other LLMs and humans.
Fine-tuning leads to near-perfect self-recognition. GPT-3.5 and Llama 2 both achieve over accuracy at self-recognition after fine-tuning on 500 examples.
Self-preference strength is linearly correlated with self-recognition. We further fine-tune LLMs to increase or decrease self-recognition, and discover a linear correlation between these two properties (Figure 1).
Definition and Measurement of Self-Preference and Self-Recognition
Self-preference is the phenomenon in which an LLM favors its own outputs over texts from other LLMs and humans.
Self-recognition is the capability of an LLM to distinguish its own outputs from texts from other LLMs or by humans.
For both definitions, we follow the prosaic rather than the intentional interpretation. That is, we use the term “self” in an empirical sense, without claiming that the LLMs have any notion or representation of itself. The prosaic interpretation allows these two concepts to exist independent of one another: An LLM can prefer texts it generated without recognizing that those texts were in fact generated by itself.
In our experiments, one LLM can play up to three different roles: generator, evaluator, and authorship identifier, depending on the instructions that they are given. Statements like “an LLM evaluator is used to evaluate texts generated by itself” emphasize the fact that the generator and the evaluator use the same underlying model, although it receives different prompts and can behave differently.
We focus on text summarization, a task where self-evaluation has been used extensively in both model-based benchmarks and reward modeling (Stiennon et al., 2020; Wu et al., 2021). For human-written summaries, we sample 1,000 news articles from each of two datasets: XSUM (Narayan et al., 2018) and CNN/DailyMail (Nallapati et al., 2016) Each example includes news article and a human-written summary.
To produce and evaluate LLM-generated summaries, we use instruction-tuned LLMs: Llama-2-7b-chat (Touvron et al., 2023), GPT-3.5 (Brown et al., 2020), and GPT-4 (OpenAI, 2023). Our fine-tuning experiments include Llama 2 (our own implementation) and GPT-3.5 (via API). We sample summaries from each of the LLMs at temperature zero using simple prompts, and standardize the punctuation and initial capitalization to match the formatting of the human summaries. The following evaluations of out-of-the-box LLMs are all based on these sampled summaries. Our code for evaluating GPT-4, GPT-3.5, and Llama 2, as well as for fine-tuning Llama 2 is available at https://bit.ly/llm_self_recognition.
2 Measurements
For both properties, we measure LLMs by prompting them using one of the following two formats:
Pairwise measurement. The LLM is presented two pieces of text, one generated by the LLM itself, and another by a second LLM or a human. We ask the LLM to choose which text has higher quality for self-preference, and which one was generated by itself for self-recognition. The identity of the alternative source is not revealed to the evaluator. We compute a prediction confidence by normalizing the output probabilities of tokens associated with the two options.
To account for the LLMs’ ordering bias (Pezeshkpour & Hruschka, 2023), we prompt the LLMs twice for each example by swapping the options, and compute the average of the two confidence scores.
Individual measurement. The LLM is presented a single piece of text either generated by itself, by another LLM, or by a human. For self-recognition, we prompt the LLM evaluator to answer a yes/no question of whether a given summary is of its own generation, with the confidence score computed in the same way as pairwise measurement. For self-preference, we prompt the LLM evaluator to rate a given summary with a Likert-scale rating of one to five for a given summary, and compute the final rating as the average of the five possible scores weighted by the output probability of each number token.
To make scores comparable across individual and pairwise measurements, we further normalize the individually measured confidence and Likert-scale scores. Let’s say in a pairwise measurement, GPT-4 is evaluated for self-recognition on a summary produced by itself and another one by Llama 2. In individual measurement, we gather these two scores separately, and divide the GPT-4’s confidence on its own summary by the sum of two scores; we consider that to be GPT-4’s confidence on this pair of examples. We similarly re-normalize the Likert scale scores for self-preference.
We refer to the results of these measurements as self-recognition and self-preference scores (noting whether they were computed in the pairwise or individual setting); taking the average of these scores over a whole dataset provides the overall self-recognition and self-preference scores of an evaluator model.
3 Out-of-the-Box Self-Recognition
Figure 2 (left) visualizes the self-recognition capability of each LLM evaluator in the pairwise setting. The results appear consistent across the two datasets, so we aggregate them for clarity. Our main observations are the following:
GPT-4 is significantly more capable at self-recognition than the two weaker models.
All three LLM evaluators most easily distinguish their summaries from human-written summaries. Human is the only alternative that Llama 2 can effective distinguish itself from.
Weaker LLMs struggle to distinguish themselves from stronger LLMs: Llama 2 is completely incapable of distinguishing itself from GPT-3.5 and GPT-4, and GPT-3 struggles to distinguish itself from GPT-4.
The GPT-4 evaluator does not distinguish Llama 2 summaries from its own summaries more easily than GPT-3.5 summaries. This goes against our intuition that self-recognition scores should increase as the dissimilarity between evaluator and evaluatee increases.
Figure 2 (right) visualizes self-recognition scores measured in the individual setting. As expected, self-recognition capability drops across the board in this setting where the LLM loses the benefit of having two pieces of texts to compare and contrast. GPT-4 is the only model capable of distinguishing authors with non-trivial accuracy. Interestingly, looking at the un-normalized confidence scores (Figure 3), we see that GPT-4 is strongly biased against predicting any text as being generated by itself, regardless of the actual author; GPT-3.5 and Llama 2 show the opposite bias.
4 Out-of-the-Box Self-Preference
The pairwise setting shows evidence of self-preference even in the absence of baseline human preference data, since the self-preference scores of each pair of evaluator models adds up to more than . All models show the greatest self-preference against the human summary.
In the individual setting, we observe that Llama doesn’t give very different scores to the different sources, while GPT-4 and GPT-3.5 show signs of self-preference. With the exception of GPT-3.5 on the XSUM dataset, evaluators show a tendency to give the human summary lower scores compared to model-generated summaries and a tendency to give their generations higher scores compared to the scores given by other evaluators (Figure 4.
5 Alternative Adjustment for Ordering Bias
All three evaluators models display ordering bias. GPT-4, GPT-3.5, and Llama reverse their pairwise preferences when the ordering of options is reversed at rates of 25%, 58%, and 89% respectively, averaged across tasks and datasets (Figure 5). We account for this bias by averaging the logit-based confidence scores across the two orderings.
An alternative interpretation of the data is, for each evaluator, to discard all the results with the label “ambiguous” where its preference displayed ordering-based reversal, reporting an evaluator’s self-recognition ability and self-preference tendency as its frequency of recognizing or preferring its own summary in “unambiguous” cases (Figure 5). This method exposes differences in evaluator results between the two datasets, but supports the presence of out-of-the-box self-recognition and self-preference.
Measuring Correlation between Self-Preference and Self-Recognition
Having validated the existence of self-preference and self-recognition, we now turn to study the correlation between the two. The main research question is to understand whether they have a causal relationship; in particular, our hypothesis is that self-recognition causes self-preference—that LLMs prefer their own outputs because they recognize them. Our investigation of the hypothesis is motivated by its safety implications, which we discuss in detail in Section 5.
We finetune LLMs to alter their self-recognition capability, and measure how the strength of self-preference changes accordingly. The correlation doesn’t by itself prove the causal hypothesis—validating the causal hypothesis would require mechanistic tools that do not exist for LLMs yet (more discussion in Section 5). But we take two approaches to provide more evidence for the hypothesis:
We invalidate the inverse causal relationship that self-preference causes self-recognition (Section 3.3).
We control for confounders by fine-tuning LLMs on tasks unrelated to self-recognition (Section 3.4).
We use supervised fine-tuning on pairwise examples to alter the LLMs’ self-recognition capability. The training examples are generated in the same way as the evaluation examples we use in the pairwise measurement. Each example consists of two summaries of a given article, one by the LLM evaluator and the other by another LLM or human source. The evaluators are trained to predict the final token, representing the the correct choice among two options, using supervised learning with cross-entropy loss. To train an LLM evaluator, we use 500 training articles and collect examples from all alternative sources (two other LLMs and human). The remaining 500 articles and the associated summaries are used as the evaluation dataset. We train on the two datasets separately but evaluate both in- and out-of-domain. GPT-3.5 was fine-tuned via API, and Llama 2 using our own implementation. The Llama models are quantized to 8 bits and fine-tuned for one epoch using Adam optimization and a learning rate of .
2 Fine-Tuning Results
For both detection and comparison tasks in the pairwise setting, ordering bias leads to the average probabilities being the only quantitative distinction in cases where the model favors the second option presented for both orderings of a pairwise question.
Fine-tuning on the self-recognition task causes models to become more confident in their pairwise judgments (Figure 6). When plotting self-recognition score against self-preference, the data clump more tightly around and the high- and low-score corners.
In the pairwise setting, fine-tuning for self-recognition results in a strong linear correlation between a model’s self-recognition ability and it’s self-preference (Figure 7). We find that fine-tuning for self-recognition ability on one dataset transfers to the other.
In additional to analyzing the relationship between self-recognition ability and overall dataset self-preference, we measure the correlation between these two properties on the example level (Table 1). For GPT-3.5 on the XSUM dataset, the evaluator prior to fine-tuning has a correlation of (Kendall’s ) between correctly recognizing its summary from a pair and preferring its summary from that same pair. Every fine-tuning configuration we employ results in a model with a positive correlation on this metric and this correlation does not change meaningfully as self-recognition ability increases.
3 Invalidating the Inverse Causal Relationship
It’s important to ensure that the causal direction is not the reverse—that the LLM does not recognize its own outputs because the quality is objectively higher, in which case the correlation does not reflect favoritism and should not raise any safety concerns. To show that the inverse causal relationship is not the case, we show that LLM evaluators do not systematically favor the summaries generated by fine-tuned models to those generated by the original model.
Some fine-tuning runs resulted in degraded generation quality (e.g., summaries that include only the characters “1” and “2”). The remainder show an average preference for the fine-tuned model’s generations of 0.46, reflecting a slight preference against the new generations. Only 22% of the evaluator-domain pairs show a preference greater than for the fine-tuned model’s generations.
4 Controlling for Confounders
To confirm that the fine-tuning procedure did not introduce any confounding effect on the variables under study, we additionally fine-tune GPT-3.5 and Llama on three tasks unrelated to self-recognition: length, vowel count, and Flesh-Kincaid readability score (Flesch, 1948). We also fine-tune both models to always output the answer “1” and to answer randomly. Together with the three fine-tuned models trained on the self-recognition task, this gives eight fine-tuned models for each of the two datasets. We find that these models can be placed in the same linear trend that we see in the models fine-tuned on the self-recognition task (Figure 1).
5 Labeling the Source
An additional experiment using only prompting to show evidence of a causal link between self-recognition and self-preference is to add either correct or reversed labels indicating the source of each summary in the pairwise self-preference question (Table 6). The GPT-4 and GPT-3.5 evaluator models show a reversal in self-preference when the labels are reversed in the XSUM dataset; that is, they show a preference for the summary labeled as being their own generation. In the CNN dataset or the Llama evaluator, no clear difference is shown from label reversal.
Related Work
The general tendency of LLMs to prefer their own generations was first recognized in the context of LLM-based benchmarks (Bitton et al., 2023a; Zheng et al., 2024; Bai et al., 2024). Liu et al. (2023), similarly to us, study self-preference bias between BERT, T5, and GPT-3.5 on text-summarization. As we discussed The larger capability gap between these models compared to ours make it difficult to control for summarization quality.
Koo et al. (2023) include self-preference in a suite of tests for LLM cognitive biases in a question-answering setting using pairwise measurements. They find GPT-4 to demonstrate lower self-preference than GPT-3.5 out-of-the-box, contrary to our findings, which suggests that evaluation on more datasets is necessary to draw a generalizable conclusion. Neither of these previous works attempted to provide an explanation for self-preference, nor did they study methods to alter self-preference strength.
2 Self-recognition and situational awareness
Hoelscher-Obermaier et al. (2023) evaluate GPT-3.5, GPT-4, and Claude-2 on their out-of-the-box self-recognition capabilities. The authors use pairwise measurement on pairs of two ten-sentence fables based on BIG-bench (Srivastava et al., 2023). On this task, contrary to our findings, GPT-3.5 is more accurate than GPT-4, which is less than 50% accurate, again showing the need to experiment on more datasets for generalizable conclusions.
Self-recognition can be seen as a form of situational awareness or self-awareness (Laine et al., 2023; Wang et al., 2024; Berglund et al., 2023; Perez & Long, 2023). Among existing work in this direction, self-recognition is most similar to calibration (Kadavath et al., 2022; Yin et al., 2023; Amayuelas et al., 2023). In this line of work, an LLM is considered to possess self-knowledge if its verbalized uncertainty is well-calibrated. Whereas they measure the correlation between verbalized uncertainty and accuracy, we measure the correlation between the uncertainty on two different tasks.
3 LLM detection
Detection of LLM-generated text is important to both AI safety and combating misinformation (Jawahar et al., 2020; Crothers et al., 2023; Wu et al., 2023; Yang et al., 2023; Kumarage et al., 2024). Despite having similar goals, self-recognition focuses on the introspective capability of language models, rather than how well a third party can discern various sources of text. The self-recognition task can be seen as a highly restricted version of detection where the method is limited to prompting the LLM. In particular, the detector LLM is not given explicit access to information such as perplexity, which in many detection methods is a crucial component (Mitchell et al., 2023; Hans et al., 2024).
Limitations, Discussion, and Conclusion
Self-recognition is a general capability that can potentially affect many multi-LLM interactions. In this paper, we focus on self-preference as the downstream property and provide initial evidence towards their causal relationship, but we see evidence that both the capability and causal hypothesis can generalize to more downstream properties. In particular, by evaluating LLMs on datasets with distinct construction processes, we observe that self-recognition fine-tuning generalizes across the two datasets and that our hypothesis holds out-of-distribution. Motivated by these results, we discuss safety risks caused by self-recognition as a general capability as well as its causal effect on various biases.
Biased self-evaluation directly affects model-based benchmarks (Shashidhar et al., 2023; Zeng et al., 2023; Yuan et al., 2023; Fu et al., 2023; Li et al., 2024): a model’s rating can be inflated simply because it’s most similar to the model used for evaluation. The bias is also a risk for methods designed for safety and alignment, such as reward modeling (Leike et al., 2018; Stiennon et al., 2020) and constitutional AI (Bai et al., 2022), for similar reasons: the reward model gives higher scores to models similar to itself, leading to weaker oversight and supervision. Such bias can be further amplified if the model is updated with feedback or training signal generated by itself (Pan et al., 2024; Xu et al., 2024).
Our work provides a basis for countermeasures against self-preference. If future evaluation confirms self-preference to be as pervasive as other biases such as ordering bias, countermeasures such as authorship obfuscation should be incorporated into standard prompting practice.
White-box adversarial attacks for free and unbounded reward hacking. In an adversarial setting (see Raina et al. (2024) for example), an LLM defender is no longer protected by black-box access if the adversary LLM recognizes their similarities. In the worst case scenario where the adversary use the same LLM as the defender, the adversary can gain unbounded access to the defender. A similar concern applies to the non-adversarial setting, where similar LLMs are use as both optimizer and reward model, as well: the strength of potential reward hacking is unbounded even if the two LLMs only communicate textually. For example, the optimizer can ignore the feedback provided by the reward model, and instead directly optimize for the shared, unaligned representation of the human-specified objectives.
2 Limitations and Future Work
Validation of the causal hypothesis. Despite the use of a diverse set of control tasks, our experiments can only provide evidence towards the causal hypothesis without fully validating it. The argument can be further strengthened by experimenting (and rejecting) more hypothesis for potential confounders between the two properties, but only up to a point—for hypotheses based on properties that we would consider to be valid explanations for self-recognition, failure to reject those hypotheses not invalidate our claim. For example, if an LLM uses verbosity as a cue to recognize itself, then the fact that fine-tuning on verbosity prediction leads to high self-recognition and self-preference doesn’t mean it’s a confounder between the two.
Controlling for groundtruth generation quality. Self-preference can be justifiable if the LLM’s generation is indeed higher quality than the alternative. From a safety perspective, what we are interested in is disproportionate self-preference, e.g., an LLM preferring its own generation even when it’s equal or worse quality than the alternative. This would require controlling for generation quality when measuring self-preference using groundtruth annotation. Although our hypothesis does not require self-preference to be disproportionate, the addition of control would improve its relevance to safety. Our existing results provide indirect evidence for disproportionate self-preference: the sum of self-preference scores of a pair of LLMs exceeds one, which means that for at least a portion of the dataset they both prefer themselves.
Example-level causal hypothesis. Our central hypothesis can be interpreted on either the example or capability level. We focus on the capability level: high self-recognition capability causes LLMs to show stronger self-preference. The example level counterpart would be: an LLM shows preference towards a piece of text because it recognizes the text as its own generation, an hypothesis of interest to interpretability. Although we observe on the correlation of the two properties on the confidence of individual predictions, our control experiments cannot further the causal argument on the example level. One approach to gather evidence for the example-level causal hypothesis is to perturb or paraphrase LLM-generated text to inhibit self-recognition and measure self-preference. Note that this task—paraphrasing text to inhibit self-recognition—has the same goal as methods designed to bypass LLM detection, so we can repurpose these methods for our experiment.
Limited number of experiment conditions. We focus on text summarization as a realistic problem with existing high quality data that have seen successful application of self-evaluation. Our cross-dataset evaluation provides initial evidence that self-recognition is a general capability that can be amplified easily by fine-tuning on a small number of examples from one dataset. Our future work will validate the hypothesis on more text summarization datasets, more tasks, as well as more frontier LLMs. We will also experiment with fine-tuning for self-recognition on the general domain rather than on a specific task.
Variance reduction. Our preliminary experiments indicate that the strength of both properties are insensitive to prompts, so all conditions use the same straightforward prompt design. To reduce variance, we will expand our experiments with more prompt designs in future work, including instructions to condition LLMs for better calibration (and reduce rejection responses). Along the lines of fine-tuning on the general domain, we will also mix self-recognition with standard instruction following datasets to improve coverage on the spectrum of self-recognition signal strength.
3 Conclusions
We provide initial evidence towards the hypothesis that LLMs prefer their own generations because they recognize themselves. In addition to evaluating LLMs out-of-the-box, we show that fine-tuning on a small number of examples elicit strong, generalizable self-recognitiono capability on summarization datasets. By varying fine-tuning task, we observe a linear correlation between self-recognition and self-preference, and validate that the correlation cannot be explained away by potential confounders. Our results establish self-recognition as a crucial factor in unbiased self-evaluation as well as an important safety-related property. The experiment design also provides a blueprint to explore the effects of self-recognition on other downstream properties.
Acknowledgements
This project has benefited from financial support to SB by Eric and Wendy Schmidt (made by recommendation of the Schmidt Futures program) and Open Philanthropy, and from in-kind support by the NYU High-Performance Computing Center and Google Cloud. This material is based upon work supported by the National Science Foundation under Grant Nos. 1850208, 1922658 and 2046556. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation. This work used the Delta GPU system at the National Center for Supercomputing Applications through allocation CIS230057 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, which is supported by National Science Foundation Grant Nos. #2138259, #2138286, #2138307, #2137603, and #2138296.
References
Appendix A Generating Summaries
We generate summaries using GPT-4, GPT-3.5, and Llama-2-7b (Table 3). We remove initial text like “Here are some highlights from the article.” For the CNN summaries, we also clean the LLM output to match the formatting of the human summaries (Table 2) by stripping bullet points or numbers from the list and removing trailing punctuation.
Appendix B Fine-Tuning on Control Tasks
Appendix C Pairwise-Setting Experiments
Prompts for the pairwise setting are shown in Table 5. For the experiments in which the summaries were labeled with either correct or incorrect sources (Section 3.5), the “Summary1” and “Summary2” portions of the prompt were followed with parenthetical “ ({source}’s summary)” to indicate the summary’s source. Table 6 shows the full results of the labeling experiments.