Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Jacob Eisenstein, Chirag Nagpal, Alekh Agarwal, Ahmad Beirami, Alex D'Amour, DJ Dvijotham, Adam Fisch, Katherine Heller, Stephen Pfohl, Deepak Ramachandran, Peter Shaw, Jonathan Berant
Introduction
To align machine learning systems with human preferences, it is common to use reward models that are finetuned on preference annotations to score potential outputs by how likely they are to be preferred by human raters (Christiano et al., 2017; Stiennon et al., 2020; Bai et al., 2022; Roit et al., 2023). There are many ways to use reward models to align policy models: they can act as training signals in reinforcement learning (Christiano et al., 2017; Stiennon et al., 2020), they can select examples for further imitation learning (Gulcehre et al., 2023; Liu et al., 2023; Dong et al., 2023; Touvron et al., 2023), or they can be applied at inference time to steer the output distribution toward higher expected reward (e.g., Yang & Klein, 2021; Gao et al., 2023). Such procedures create a semi-adversarial dynamic in which the language model is encouraged to produce outputs that obtain high reward by exploiting errors in the reward model. Furthermore, while the reward model is trained on a fixed set of human preference data, the process of alignment shifts the distribution of its inputs, increasing the likelihood of such errors. This phenomenon where the policy language model exploits reward model errors is often termed reward hacking (Amodei et al., 2016), reward gaming (Skalse et al., 2022; Pang et al., 2023), or reward over-optimization (Gao et al., 2023).
Reward hacking has been investigated from several perspectives in prior work (e.g., Krakovna et al., 2020; Skalse et al., 2022; Pan et al., 2022). Bai et al. (2022) used reinforcement learning with human feedback (RLHF) and trained two reward models on non-overlapping splits of preference data, using one to drive alignment, and the other to measure the quality of the outputs. They find that RLHF increases performance according to both the driver and measurement models, but that a performance gap emerges as the policy is allowed to diverge from the initial distribution. However, both reward models were built on base models trained on the same pretraining data, which, as we will show, limits their diversity (as hypothesized by Gleave & Irving (2022)) and thus may understate the effect of reward hacking. Other work has simulated the relationship between a “true” reward and a learned proxy, showing that it is possible to over-optimize the proxy to such an extent that the true reward starts to decrease (Gao et al., 2023; Coste et al., 2023). This has been replicated in more realistic settings by examining (and creating) spurious correlations in reward model training data (Pang et al., 2023).
In this work, we first analyze reward model distribution shift from the perspective of underspecification (D’Amour et al., 2022), which occurs when a machine learning pipeline yields reliable performance on held-out data from the training distribution, but variable performance on out-of-distribution data. When applied to learning reward models from human preference data, we show that reward models that agree in-distribution often disagree when transferred out-of-distribution. Furthermore, such disagreements are more pronounced when the reward models are built on different pretrainings, even when that difference is induced merely by varying the pretraining random seed. These disagreements become increasingly severe when evaluated on outputs of a policy model that has been aligned to a specific reward model. This occurs both when using reward models in RLHF, as well as when using an inference-time alignment procedure, best-of- reranking, where samples are drawn from the policy and then reranked with a reward model.
Motivated by these findings, we systematically investigate reward model ensembles as a possible remedy for reward hacking. Assuming different models err in different ways, ensembling can leverage reward uncertainty across the ensemble during alignment (see Figure 1, Left). We explore several techniques for aggregating scores across the ensemble, e.g., taking the median score as a robust estimate of the true reward of the policy. We also consider two types of ensembles: pretrain ensembles, where different members of the ensemble differ in the random seed used during the pretraining phase, and finetune ensembles, where members differ only in the random seed used during finetuning. These ensembles are then evaluated across several types of policies and preference annotations: dialogue preferences for a helpful assistant (Bai et al., 2022), summarization quality (Stiennon et al., 2020), and whether a single-document summary is grounded in its source text (Roit et al., 2023).
We find that pretrain ensembles substantially outperform finetune ensembles. Moreover, they consistently outperform single reward models, unlike finetune ensembles, which in many cases are comparable to single reward models. However, our analysis also reveals that policies trained with ensembles are still susceptible to reward hacking: different reward models sometimes share similar error patterns, which in turn propagate to the ensemble (see Figure 1, Right). This is exploited and amplified by the policy, leading, for example, to outputs that are too short when tuning for factuality, too verbose when tuning for summarization quality, or responses that follow a particular format that is often unsuitable, when training a helpful assistant. Thus, it is possible that methods that, unlike ensembles, are aware of the distance of outputs from the reward data distribution (Liu et al., 2020) could provide more reliable estimates of uncertainty.
In concurrent work, Coste et al. (2023) argue that reward model ensembles effectively mitigate reward hacking. Our work shares a similar research question, but differs in several ways, leading to more nuanced conclusions. First, we investigate the difference between pretrain and finetune ensembles, finding that pretrain ensembles are considerably more effective. Second, we use human-annotated preference data rather than synthetically-generated labels, which provides a more realistic experimental setup. Third, we perform analysis that demonstrates the limitations of reward ensembles, showing reward ensembles are still susceptible to reward hacking. Last, our experimental setup covers a wider range of tasks, larger reward models, and more extensive policy optimization.
Preliminaries
Reward models have become the primary tool for aligning LMs towards user-facing applications. We now briefly review how reward models are trained (§2.1) and how they are used for alignment (§2.2). We then describe the experimental setup that we will use for the remainder of the paper (§2.3).
We focus on the the typical setup where reward models are trained from preference data, , where is annotated to be preferred over for prompt . Under the Bradley-Terry model (Bradley & Terry, 1952), the probability that response is preferred over given a reward function and a prompt is , where is the sigmoid function. Then, we can use preference data to train a reward model by maximizing
The Bradley-Terry model is underdetermined: for any reward model , we can define an equivalent reward model, where is a prompt-dependent constant, obtaining the same objective value as , i.e., . This is problematic for ensembling: if different reward models choose different values for , then order statistics like median and minimum are meaningless. We therefore modify the objective function by adding a regularization term to encourage the sum of reward values per preference pair to stay close to zero, i.e.,
where is a small positive value, thereby resolving the issue of underdetermination.
Note that reward models can also be trained from “pointwise” data, such as toxicity or factuality annotations on individual examples (Yang & Klein, 2021; Roit et al., 2023). Such reward models are not underdetermined and so can be aggregated without adjustment.
2 Aligning Language Models using Reward Models
Best-of- reranking (BoN) is an inference-time alignment strategy, where given a prompt , we sample generations from a policy language model and return the generation that has the highest reward according to a reward model , i.e., . The Kullback–Leibler (KL) divergence of BoN from the initial policy is upper bounded by . BoN tends to outperform more elaborate alignment techniques like RLHF in the low-KL regime (Gao et al., 2023), albeit with the cost of generating multiple samples at inference time.
Reinforcement Learning from Human Feedback (RLHF) is an online reinforcement learning method that trains a policy language model to maximize expected reward, while staying close to an initial policy, , which is typically finetuned on supervised data (prompt-output pairs). Distance from the initial policy is measured with KL divergence, which leads to the regularized objective
where is a reward model, is a distribution over prompts, and is a hyper-parameter. Typically, this objective is optimized using PPO (Schulman et al., 2017), which we also use in this work.
3 Experimental Setup
We will examine the performance of reward models (both single models and ensembles) across three tasks. An example from each task is provided in Table 1.
tl;dr: A summarization benchmark where authors summarize their own reddit posts (Völske et al., 2017). We use the preference data created by Stiennon et al. (2020). This benchmark has been commonly used to evaluate finetuning of policy LMs (Rafailov et al., 2023; Zhao et al., 2023).
helpfulness: A helpful assistant benchmark (Bai et al., 2022), where given a partial conversation between a human and a digital assistant the goal is to complete the next turn of the assistant. This benchmark has also been commonly used for evaluating finetuned policy LMs (Bai et al., 2022; Rafailov et al., 2023). We use the base dataset (44K examples), where responses are generated from a 52B context-distilled LM, and split the training set into two: half for training the reward model, and half for training the policy model.
xsum/nli: We adopt the setup of factually-consistent summarization (Roit et al., 2023), where a model trained on XSum (Narayan et al., 2018) is finetuned to generate summaries that are consistent with the source document according to a Natural Language Inference (NLI) reward model.
Training reward models
To examine the effect of pretraining on reward models, we pretrain five T5 models from scratch with the base (220M parameters), large (770M), and XL (3B) architectures, using the standard denoising objective over the C4 corpus (Raffel et al., 2020). The pretrained checkpoints differ only in their random seed, which controls parameter initialization and the sample from the pretraining data. The same pretrained models are used for finetuning across all tasks.
We finetune each pretrained model five times using different random seeds across all three benchmarks. In tl;dr and helpfulness we use the aforementioned preference data. For xsum/nli, we finetune NLI models on the ANLI dataset (Nie et al., 2020). Overall we obtain 25 reward models per task (5 pretrain 5 finetune). This makes it possible to evaluate the effect of pretraining and finetuning on underspecfication (§3) by constructing ensembles that differ in either pretrain or finetune seed (§4).
Alignment strategy
We use the publicly available T5-large model (Raffel et al., 2020) as a policy for the two summarization tasks. For helpfulness, the task requires substantial background knowledge, and thus we use the instruction-tuned PALM-2-XXS model (Anil et al., 2023). Prior to alignment, we create a finetuned policy by finetuning on supervised data in the standard manner. We finetune on annotated summaries from tl;dr and xsum/nli for the corresponding tasks, and on the preferred responses, , from the preference data in helpfulness.
Evaluation
We use two metrics to quantify generalization of reward models—reward by a larger model and win rate. Similar to past work (Gao et al., 2023; Coste et al., 2023), we use a larger reward model to evaluate the generalization of models trained with a smaller reward model. We train a T5-XXL reward model by taking the publicly available T5-XXL (Raffel et al., 2020) and finetuning it as described above. Table 2 details the performance of reward models of different sizes on the three tasks, and it can be seen that T5-XXL outperforms the best T5-XL model. We report both average reward of the T5-XXL evaluator as well as win rate, which is the fraction of prompts for which the response sampled from the aligned policy has higher reward compared to .
The errors of the T5-XXL autoeval model might correlate with errors of the smaller T5 models because they are trained on the same preference data. For this reason, we also evaluate win rate according to a prompted PaLM-2-Large model, which was not exposed to the reward training data but was instruction-tuned on FLAN (Wei et al., 2022). Given a prompt , we sample a response from and from . We then ask PaLM-2 which response is better, using a hand-engineered prompt proposed by Rafailov et al. (2023). To avoid position bias we run PaLM-2 on the two possible orderings and , sample outputs for each order and determine the winner on this prompt through majority voting. This style of evaluation has become common recently (Dubois et al., 2023; Singhal et al., 2023) and was shown to correlate well with human judgements (Rafailov et al., 2023).
Underspecification in Reward Models
We now analyze alignment strategies that use a single reward model, and demonstrate that reward models are underspecified. First, Table 2 shows the average in-distribution accuracy across the 25 different reward models, together with the standard deviation (which is low in-distribution).
The story changes, however, when we move to out-of-distribution data. Figure 2 shows the expected reward achieved by BoN as a function of the number of sampled candidates, , for three reward model scales (KL is approximately ). The dotted green line shows the expected reward of the top-ranked output according to the reranker itself, while the dashed orange line shows the expected reward of the same output according to reward models that share a pretrain seed. The solid blue line shows the expected reward according to reward models that do not share a pretrain seed. Unsurprisingly, the reranker scores its own top outputs more favorably than the other reward models do. However, the reranker’s outputs are scored significantly less favorably by reward models which do not share a pretrain with the ranker. Reward models that share a pretrain seed with the ranker model overestimate the true reward of the top-ranked output—suggesting that finetune ensembles are not sufficiently diverse because of the shared pretraining state of each of the ensemble’s members. Notably, this gap does not disappear with scale, and is present for base, large, and XL models.
Moving to alignment, differences in estimated rewards induce different policies from the BoN strategy: Figure 3 shows the effects on agreement of the top-ranked summary when reward models do (crosses) or do not (circles) share pretraining seeds. Different reward models tend to produce different 1-best outputs. Again these differences are strongly associated with the pretraining seed: for example, two reward models from different pretrains will choose a different best-of-16 output more than half the time for both tl;dr and helpfulness and in all scales.
Last, Figure 4 analyzes the evolution of agreement of the estimated reward scores when performing RLHF on tl;dr for reward models of various scales. Specifically, we align a policy using a single reward model, and then measure how well pairs of reward models agree on the ranking of samples from that policy using Spearman rank correlation. To compute Spearman, we sample 5 completions for each prompt in the validation set from a policy model, at 2K step intervals during RLHF. We compare the agreement between a set of 5 reward models that share the same pre-training seed and a set of 5 that do not (both sets include the reward model used to drive RLHF). For each prompt, we compute Spearman correlation across all ten pairs in each set and report the mean correlation over the pairs. The correlation of models that do not share a pretrain is lower compared to models that share a pretrain seed. Moreover, correlation goes down during RLHF, indicating that the uncertainty about the true reward increases as a result of alignment.
Overall, our analysis demonstrates that (1) different reward models tend to disagree on out-of-distribution data, particularly when the reward models have different pretraining seeds; (2) this propagates to the trained policy model, in the sense that the resulting policy is highly tuned to the preferences of the specific reward model used to drive it; and (3) as a result, the disagreement between reward models tends to increase during alignment. These findings suggest that reward model ensembles might mitigate reward hacking, which we turn to next.
Reward Model Ensembles
We describe how to construct reward model ensembles (§4.1), and evaluate their performance (§4.2).
We showed that reward models are underspecified—as they are used more in alignment, they induce a stronger distribution shift in the outputs of the policy, which in turns leads to higher disagreement across reward models. Thus, a natural mitigation strategy is to ensemble multiple reward models, under the assumption that different models will have different errors. Aggregating over the scores of the ensemble members will help when some of the ensemble members erroneously assign high reward to a bad output.
Given a set of reward models , we define the reward of the ensemble to be , with agg indicating an aggregation function (Dietterich, 2000; Lakshminarayanan et al., 2017; Raffel et al., 2020; Zaidi et al., 2021). Intuitively, the aggregation function should be conservative, and return a lower score when there is disagreement between the ensemble members. We consider the following simple aggregation function: mean, median, and mean_minus_std, which subtracts the standard deviation of the reward from the mean to penalize high variance. We also experiment with min, but overall find it to be inferior to the alternatives.
We evaluate two types of reward ensembles: pretrain ensembles, where each member was pretrained using a different random seed,Pretraining does not complete a single epoch over the pretraining data, and thus the data observed by each member of a pretrain ensemble is different (but sampled from the same distribution). and finetune ensembles, where all members share the same pretraining seed, but use a different seed when finetuned on the reward data (which typically includes preference pairs, where one output is preferred over another). In all cases the ensemble contains exactly 5 individual reward models. Pretrain ensembles are significantly more expensive to train, but are more diverse and hence likely to lead to a more robust reward estimate. In fact, Gleave & Irving (2022) reported negative results when using reward ensembles and hypothesized this is due to ensemble members sharing the same underlying pretrained model.
2 Experiments
We now evaluate reward model ensembles across all tasks. Figure 5 shows the results of ensembling in best-of- reranking, as measured by an XXL-scale fine-tuned reward model. Pretrain ensembles consistently improve performance over individual reward models, especially for higher values of for both tl;dr and helpfulness. Finetune ensembles, conversely, improve performance in some cases and are comparable in others. For example, on tl;dr a pretrain ensemble with the mean aggregator achieves a win rate of 90% over the SFT outputs at the XL scale, while the win rate of a finetune ensemble with the same mean aggregator is 87.3%. The win rate of the average individual XL-scale reward model is 85.3% (see Table 7). For visual clarity, in Figure 5 we show only two aggregators: mean and mean_minus_std; see Appendix A for results with other aggregators. In general, the differences between aggregators are small, with mean usually performing at, or near, the top. More conservative aggregators (min and mean_minus_std) come out slightly ahead of mean at the smaller scales on tl;dr, suggesting that high variance may be a bigger issue in this setting.
Figure 6 shows the KL-reward trade-off of ensemble reward models in RLHF for tl;dr and helpfulness (evaluated with the finetuned T5-XXL model). In such plots, a better model is one that improves reward and/or reduces the value of KL from the original SFT policy (Gao et al., 2023; Coste et al., 2023). Indeed, similar to BoN, pretrain ensembles consistently outperform both finetune ensembles as well as the average individual model. We present results for the median and mean aggregators for visual clarity, and report full numerical results in Appendix B. In RLHF, KL values are much higher than BoN (which is bounded by for ). Consequently, in this setting we witness explicit reward hacking, in which the T5-XXL rewards decrease even as the RLHF objective improves. This happens most prominently for individual models, in many cases for finetune ensembles, and most rarely for pretrain ensembles—where T5-XXL reward scores decrease only when RLHF uses a T5-Base reward model. Thus, our experiments on real data yield more negative conclusions than Coste et al. (2023) about the potential of ensembles to eliminate reward overoptimization.
Because the T5-XXL autoeval model is trained on the same data distribution as the reward models used for best-of- and RLHF, it may overstate their performance. For this reason, we also use a zero-shot autoeval model (PaLM-2-Large), as described in Section 2.3. Because this evaluation is more computationally expensive, we apply it only to the largest-scale reward models (XL). Results are shown in Figure 7. Ensemble reward models consistently achieve higher win rates on both tasks and with both alignment techniques. For best-of-, pretrain ensembles get significantly higher win rates on tl;dr at ( by a permutation test); on helpfulness the differences between ensembling techniques are not significant at . On both tasks, single reward models are significantly worse, . For RLHF, pretrain ensembles generally achieve better or equal win rates at lower KL divergence from the reference policy, with particularly strong performance on helpfulness. Overall, these results mirror the T5-XXL evaluation, with one interesting difference: the PaLM-2 autoeval model reveals more reward hacking for RLHF, where win rate decreases with KL. This suggests that fine-tuned autoevaluators can overestimate performance when they are trained on the same preference data as the alignment reward models.
Figure 8 shows RLHF results for xsum/nli. Here we see a relatively small improvement for ensembles compared to individual models, and a very small difference between pretrain and finetune ensembles. We conjecture this is because xsum/nli optimizes for a particular aspect of the response, namely its factuality. This allows all models to find simple and similar strategies that lead to high reward (for example, emitting short responses with limited content), and thus ensembling does not lead to large gains in performance. We further elaborate on this when discussing limitations of ensembles in §5.
When do Reward Model Ensembles Fail?
We saw that ensembles improve performance according to automatic evaluation metrics. We now conduct a complementary analysis that illustrates that, for some types of errors, ensembling is ineffective. When all reward models share a similar error pattern, this error propagates to the ensemble. Systematic errors across ensemble members can arise due to biases in the finite reward model training data.
To demonstrate this, we manually analyze ensemble outputs to detect frequent errors, and then perform a qualitative analysis. Figure 9 shows the results of this analysis on all three benchmarks. The x-axis corresponds to outputs of the model after training for a certain number of steps, and the y-axis is a statistic of interest (e.g., average output length). We plot the statistic value for the pretrained ensemble (using mean as a representative aggregation function) and for its members. In addition, for tl;dr and helpfulness, where the reward model is trained on the preference data, we show the statistic value on the preference data validation set, conditioned on the label ‘Preferred’ or ‘Rejected’.
For helpfulness (Figure 9(a)), outputs tend to be in a format of a list, and thus we write a regular expression that captures this format. The fraction of outputs that have this pattern increases to roughly 50% for 3 members of the ensemble and to the ensemble itself. Looking at the preference data, we do not detect a tendency to produce list outputs in the preferred responses, as the fraction of outputs that matches this format is roughly 8% for both the preferred and rejected responses.
For tl;dr (Figure 9(b)), RLHF alignment leads to longer summaries (Singhal et al., 2023) and also outputs that are more extractive, i.e., copy more from the input. Summary length in characters grows substantially for the ensemble and all its members, where for the ensemble, length increases by a factor of two. On the preference data, indeed preferred responses are slightly longer than rejected responses, but much shorter than outputs post-RLHF. We also compute the longest common subsequence (in characters) between the document and the summary and find that it increases for the ensemble from 28.2 to 49.1. Again, the tendency for copying from the document already occurs in the preference data to a small degree, but is amplified by RLHF.The distribution of outputs in the preference data is not identical to the distribution of outputs before RLHF, and therefore the statistics after zero training steps do not necessarily match those of the preference data.
For xsum/nli (Figure 9(c)), training for factuality tends to make summaries shorter. Additionally, precise numbers are typically omitted from the summaries. Figure 9 shows how all members of the ensemble and the ensemble itself exhibit this phenomenon, with length in characters decreasing rapidly, as well as the fraction of examples that contain any numeric value whatsoever.
Overall, these qualitative findings are symptoms of the tendency for different pretrain reward models to learn to associate certain features with high reward. Policy models can then exploit this association, and use these features to produce outputs that are dramatically different from the reward training data, and that achieve (spuriously) high reward for both single reward models and the ensemble.
Why does this happen for both single reward models and reward model ensembles? As one indication, Lakshminarayanan et al. (2017) have proposed distance-awareness, i.e., the ability to quantify the distance of an example from the training set, as a necessary condition for achieving good uncertainty estimates. They showed in a synthetic binary classfication setup that deep ensembles provide good estimates when examples are on the decision boundary, but underestimate uncertainty in areas that are far from the training distribution. In LM alignment, the policy can shift the output distribution away from the decision boundary to areas where all reward models erroneously extrapolate in the same manner. While we focus on ensembles in this work, we hypothesize that the same phenomenon will occur in other approaches for uncertainty estimation that are not distance-aware, such as Monte-Carlo Dropout (Gal & Ghahramani, 2016) and Epistemic Neural Networks (Osband et al., 2021).
Conclusion
In this work, we investigate reward model ensembles as a method for mitigating reward hacking. We find that diversity of the reward ensemble is crucial, and that a pretrain ensemble that contains members that do not share a pretrain seed leads to stronger generalization during alignment when compared to an ensemble whose members share a pretrain seed. However, reward ensembles are not always effective—for example, we find that they can still assign reward based on spurious correlations between the input and the label. If all members of the ensemble capture the same correlations, the ensemble will inherit the same undesirable behaviour. In such cases, the policy can exploit this vulnerability and shift the distribution towards outputs that overuse this correlation, which results in reward hacking. Consequently, reward model ensembles mitigate, but do not fully eliminate, reward hacking. Future work should examine methods for uncertainty estimation that are more robust to the type of distribution shift that occurs during alignment, particularly those that are aware of how different model policy outputs are from the preference data—such as Gaussian processes (Kuss & Rasmussen, 2003; Chu & Ghahramani, 2005; Liu et al., 2020) and conformal prediction under covariate shift (Tibshirani et al., 2019).
Thanks to Sharat Chikkerur, Mohammad Havaei, and the anonymous reviewers for feedback on this paper. The research also benefited from feedback from David Bruns-Smith, Ming-Wei Chang, Michael Collins, Patrick Fernandez, Mandar Joshi, Rishabh Joshi, Balaji Lakshminarayanan, Kenton Lee, Kristina Toutanova, Victor Veitch, and Zihao Wang. Finally, we thank the people who built the infrastructure used in our experiments, including the T5X team and Léonard Hussenot, Johan Ferret, Robert Dadashi, Geoffrey Cideron, Alexis Jacq, Sabela Ramos, Piotr Stanczyk, Sertan Girgin, Danila Sinopalnikov, Amélie Héliou, Bobak Shahriari, Bilal Piot, Matt Hoffmann, Nikola Momchev, and Olivier Bachem.
References
Appendix A Numerical results for best-of-n reranking
Average agreement between reward models are shown in Tables 3-6. Autoevaluation results are shown in Table 7 and 8.
Appendix B Numerical results for RLHF
Full RLHF numerical results for helpfulness and tl;dr are shown in Table 9 and Table 10.
Appendix C Hyperparameters
We provide the hyperparameters for reward model and RLHF training in Table 11 and Table 12. For reward models, we use the validation set to choose the best checkpoint along training. For RLHF, we take the last checkpoint.