RewardBench: Evaluating Reward Models for Language Modeling
Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A. Smith, Hannaneh Hajishirzi
Introduction
Reinforcement learning from human feedback (RLHF) is a necessary but largely non-reproduced tool underlying the success of popular large language models (LLMs) such as OpenAI’s ChatGPT (Schulman et al., 2022) and Anthropic’s Claude (Bai et al., 2022a). The prevalence of RLHF stems from its efficacy at circumventing one of the greatest difficulties in integrating human values and preferences into language models: specifying an explicit reward (Christiano et al., 2017). Reward models (RMs) are central to this process. They are created by taking copies of the original language model and training those copies on labeled preference data, producing a model that can predict whether a user is likely to prefer one piece of text over another. A reinforcement learning optimizer then uses this reward model signal to update the parameters of the original model, improving performance on a variety of tasks (Ouyang et al., 2022; Bai et al., 2022a; Touvron et al., 2023).
While the post-RLHF model (known as the policy) and even the pretrained model are extensively documented and evaluated, the basic properties of the RLHF process receive far less attention. Though reward models are central to understanding the effectiveness of RLHF and moreover provide a potential glimpse at how human values map onto language models, they remain catatonically under-evaluated. Recent work on training reward models (Zhu et al., 2023a; Jiang et al., 2023c) has begun to fill this gap, but utilizes validation sets from previous RLHF training processes, such as Anthropic’s Helpful and Harmless data (Bai et al., 2022a) or OpenAI’s Learning to Summarize (Stiennon et al., 2020), which are known to have ceilings on accuracy between 60 and 70% due to inter-annotator disagreement (Wang et al., 2024). Similar investigations have yet to be conducted for Direct Policy Optimization (DPO) models. Moreover, newly released preference data aiming to expand the diversity of preference training datasets such as UltraFeedback (Cui et al., 2023) and Nectar (Zhu et al., 2023a), do not have test sets, necessitating a new style of evaluation for RMs.
We begin to rectify the lack of evaluation techniques by introducing RewardBench, the first toolkit for benchmarking reward models. RLHF is inherently a broadly applicable process. It is used to enhance specific capabilities of language models such as safety (Dai et al., 2023) or reasoning (Lightman et al., 2023; Havrilla et al., 2024a) as well as general capabilities such as instruction following (Ouyang et al., 2022) or “steerability” (Askell et al., 2021; Bai et al., 2022a). Evaluations for reward models must cover all of these categories.
In this work, we curate new data and repurpose prompts from a variety of LLM evaluation tool-kits to create structured comparisons across a variety of reward model properties. Each sample is formatted as a prompt with a manual or human-verified chosen and rejected completion. We design subsets so as to vary in difficulty. Some are constructed such that many reward models can differentiate chosen from rejected completions, reaching nearly 100% accuracy. Others are more difficult and state-of-the-art language models only reach the 60 to 70% range.
In addition to introducing a holistic benchmark, we aim to map the current landscape of openly available reward models via a reward model leaderboard. To do this, we evaluate most of the available models such those trained as classifiers, including UltraRM (Cui et al., 2023), Starling (Zhu et al., 2023a), PairRM (Jiang et al., 2023c), SteamSHP (Ethayarajh et al., 2022), models from Reward rAnked FineTuning (RAFT) (Dong et al., 2023), and others. We also evaluate popular chat models trained with Direct Preference Optimization (DPO) (Rafailov et al., 2023), for example, Zephyr- (Tunstall et al., 2023), Qwen-Chat (Bai et al., 2023), StableLM (Bellagente et al., 2024), and Tülu 2 (Ivison et al., 2023) to ground recent debates on RLHF methods and showcase specific datasets where they fall short.
With these models, we compare scaling, test reasoning capabilities, highlight three buckets of refusal behavior, and share more details on the inner workings of RMs. The accompanying code-base provides a common inference stack for many variations of models and we release many text-score pairs to analyze their performance.
Release a common framework for evaluating the many different architectures of reward models, along with tools for visualization, training, and other analysis. We also release all data used in the evaluation, composed of text-score pairs for all inputs, to enable further data analysis on the properties of reward models.Data is here: https://huggingface.co/datasets/allenai/reward-bench-results.
Illustrate the differences between DPO and classifier-based reward models across a variety of datasets. DPO models, while more plentiful due to the method’s simplicity, fail to generalize to popular preference data test sets and present a higher variance in performance.
Chart the landscape of current state-of-the-art reward models. We showcase the scaling laws, the propensity to refuse (or not), the reasoning capabilities, and more for popular RMs.
Show the limitations of existing preference data test sets for evaluating these models, showcasing common pitfalls of RMs on subtle, but challenging instruction pairs (e.g. intentionally modified rejected responses, which superficially look high quality but answer the wrong prompt).
We hope this benchmark enables more advanced reward model training, scientific understanding of the integration of human preferences in LMs, and ultimately better aligned, open language models.
Related Works
Using Reinforcement Learning to align language models with human feedback or preferences (Christiano et al., 2017; Ziegler et al., 2019) has led to improved chat models such as ChatGPT (Schulman et al., 2022) and Llama2 (Touvron et al., 2023). Incorporating human feedback into models in this way has been used to improve summarization (Stiennon et al., 2020; Wu et al., 2021), question answering (Nakano et al., 2021), image models (Lee et al., 2023) and instruction following in general (Ouyang et al., 2022).
RLHF for alignment has been operationalized beyond general preference by comparing aspect-based preference, where aspects could be more general concepts like helpfulness or harmlessness (Bai et al., 2022a), or more fine-grained ones (Wu et al., 2023), among others. In general terms, RLHF involves training a reward model on preference data collected from crowdworkers (Wang et al., 2024) (or via using an LLM as a judge of responses, denoted RL from AI Feedback (Bai et al., 2022b)). Given a reward model, a policy can be learned using RL algorithms like PPO (Schulman et al., 2017), a method that has been shown to work well for language policies (Ramamurthy et al., 2022). Another option is to directly optimize a model with chosen and rejected pairs, using DPO (Rafailov et al., 2023). Some reward modeling extensions include process reward models (Luo et al., 2023; Lightman et al., 2023) and step-wise reward models (Havrilla et al., 2024b), which are used for reasoning tasks to provide a correctness label for each of the steps in a reasoning chain.
Despite RLHF’s impressive results, the approach has also been shown to lead to overoptimization (Gao et al., 2023) and divergence from the initial data distribution (Marks et al., 2024). Such reward hacking might be partially, but not fully, mitigated using RM ensembles (Coste et al., 2023; Eisenstein et al., 2023), weight averaging (Ramé et al., 2024), or constrained optimization (Moskovitz et al., 2023).
2 Reward Model & RLHF Evaluation
Preference tuned models can be evaluated using downstream evaluations, for example using the AlpacaFarm (Dubois et al., 2024) framework. In AlpacaFarm, LLMs are used to simulate human preferences by comparing a model generated output with the output from a reference model. The reported metric is the win-rate of the model over the reference model, which is being calculated over a set of 805 instructions. Similarly, MT-Bench (Zheng et al., 2023), evaluates chatbots on multi-turn conversations that are judged by LLMs as proxy for human judgments. Chatbot Arena (Zheng et al., 2023) is an evaluation benchmark that crowdsources the preferences between two different model outputs. These types of setups do not directly evaluate the reward model.
Other works, on the other hand, analyze the reward model, such as Singhal et al. (2023), who look at the training dynamics of RMs. In their study they found a strong correlation between output length and rewards. Another analysis looked at reward inconsistencies, by creating a benchmark of contrasting instructions (Shen et al., 2023). Most importantly they found that reward model inconsistency also affects the RLHF training and resulting RLHF’ed model.
Background
The first step of training a reward model, and therefore doing RLHF, is collecting preference data from a group of human labelers. Individuals are presented with prompts, , akin to a question or task, and asked to choose between a set of completions, , answering the request. The most common case is for only two completions to be shown with measurement of preference, such as win-loss-tie or a Likert scale indicating the magnitude of preference between completions (Bai et al., 2022a), though other methods for labeling exist, such as ranking in a batch of 4 to 7 answers (Ouyang et al., 2022). The resulting data is transformed into a set of prompt-chosen-rejected trios, where the chosen completion is preferred over the rejected completion for training.
Training a reward model involves training a classifier to predict the human preference probability, , between two answers, as modeled by a Bradley-Terry model (Bradley and Terry, 1952):
Then, estimate the parameters of the reward model by optimizing the maximum likelihood loss as follows:
For language models, the RM is often implemented by appending a linear layer to predict one logit or removing the final decoding layers and replacing them with a linear layer. At inference time, a trained reward model returns a scalar, such that (which intuitively is the probability that the completion would be a preferred response, but is trained indirectly via the pairwise loss). Thus, a win between completions and is achieved when .
2 Direct Preference Optimization
Direct Preference Optimization solves the RLHF problem without needing to learn a separate reward model. It arranges an reward function from the model probabilities, directly optimizes the RM, and extracts a language model from it (Rafailov et al., 2023). The implicit reward used in DPO is a function of the policy model probabilities (i.e. the model being trained), , a regularization constant, , the base model probabilities, , and a partition function :
Given two completions to a prompt, we compare the rewards and as follows, where the score is computed via the log ratios of :
The RewardBench Benchmark
In this section, we detail the design philosophy and construction of the evaluation dataset. The dataset is designed to provide a broad set of basic evaluations for reward models, covering chat, instruction following, coding, safety, and other important metrics for fine-tuned language models. The RewardBench dataset contains a combination of existing evaluation prompt-completion pairs, and those curated for this project.
A good reward function, and therefore a good RM broadly, is one that stably assigns credit to the classes of good or bad content.There are more considerations on how to use a RM, but the initial notion of quality should be one that agrees with curated data. Next, we can evaluate which RMs are best for downstream tasks such as RLHF. Given one verified answer that is better than another for factual or clear qualitative reasons (e.g. typos), a good reward model will choose the correct one 100% of the time. To evaluate this, each datapoint consists of a prompt and two completions, chosen and rejected. For each prompt, the score of the reward model is computed. The prompt is then categorized as a win if the score of the prompt with the verified chosen completion is higher than that of the verified rejected completion, as shown in Fig. 1. Finally, we report accuracy for each subset as the percentage of wins. For all the section scores of RewardBench (e.g. Chat or Safety) except Prior Sets, the average score is weighted per-prompt in the requisite subsets.
The benchmark is broken down into five sections from different subsets – the first four compose the RewardBench dataset described in this section. We have broken down the dataset into these subsections to create one final RewardBench score in order to reasonably weigh different aspects of an RM’s performance. The summary of the dataset is shown in Tab. 1 (see appendix B for full details) At a high level, the subsets consist of the following:
Chat: Testing a reward model’s basic ability to distinguish a thorough and correct chat response in open-ended generation. Prompts and chosen, rejected pairs are selected from AlpacaEval (Li et al., 2023b) and MT Bench (Zheng et al., 2023) completions, two popular open-ended chat evaluation tools.
Chat Hard: Testing a reward model’s abilities to understand trick questions and subtly different instruction responses. Prompts and chosen, rejected pairs are selected from MT Bench examples with similar ratings and adversarial data specifically for fooling LLM-as-a-judge tools from LLMBar’s evaluation set (Zeng et al., 2023) (reformatted for RMs).
Safety: Testing the models’ tendencies to refuse dangerous content and to avoid incorrect refusals to similar trigger words. Prompts and chosen, rejected pairs are selected from custom versions of the datasets XSTest (Röttger et al., 2023), Do-Not-Answer (Wang et al., 2023), and examples from an in-development refusals dataset at AI2, where the chosen response is a refusal and the rejected is harmful text of either dangerous or offensive nature.
Reasoning: Evaluating the models code and reasoning abilities. Code prompts are created by reformatting HumanEvalPack examples with correct code as chosen and rejected as one with bugs (Muennighoff et al., 2023). Reasoning prompts pair reference answers with incorrect model generations from the PRM800k dataset (Lightman et al., 2023).
Prior Sets: For consistency with recent work on training reward models, we average performance over test sets from existing preference datasets. We use the Anthropic Helpful split (Bai et al., 2022a) (the only multi-turn data), the Anthropic HHH subset of BIG-Bench (Askell et al., 2021), a curated subset of the test set from the Stanford Human Preferences (SHP) Dataset (Ethayarajh et al., 2022), and OpenAI’s Learning to Summarize Dataset (Stiennon et al., 2020).The dataset with more test sets and details is found here: https://huggingface.co/datasets/allenai/preference-test-sets
2 RewardBench Scoring
The primary scoring metric for RewardBench is accuracy. For each prompt-chosen-rejected trio, we infer the score the reward model assigns for the prompt-chosen and prompt-rejected pairsFor some reward models, such as PairRM and SteamSHP, their intended use is with pairwise inputs, so we evaluate in that manner following the original source code. then assign a true classification label when the chosen score is higher than rejected. This technique is highlighted in Fig. 1. More details on scoring, including for DPO models, is included in Sec. 3.
Given the binary classification of correct or not, a random model achieves a result of 50 on our benchmark. On many subsets, models achieve at or well below the random baseline, indicating substantial areas of progress in reward models.
In order to create a representative, single evaluation score, we perform a limited mixture of averaging across results. For all the subsets detailed in Sec. 4.1 except for Reasoning, we perform per-prompt weighted averaging across all the prompts in the subset to get the section score to normalize by the size of each category. For example, in Chat we take a weighted average of the AlpacaEval and MT Bench sets based on the number of prompts. For Reasoning, we increase the weight of the PRM-Math subset so code and math abilities are weighed equally in the final number, rather than increasing the relevance of code. For Prior Sets, we take an unweighted average over the subsets due to the large disparity in dataset sizes. Once all subsets weighted averages are achieved, the final RewardBench score is the average across the subset scores.
Evaluation Results
RewardBench includes evaluation of many public reward models, ranging in parameter count from 400 million (PairRM) to 70 billion (Tülu 2), trained as classifiers or with Direct Preference Optimization (when the reference model is available). In this section, we detail the core findings of RewardBench and more results are available in Appendix A. In particular, we study the state-of-the-art reward models (Tab. 2), results of similar-size models at 7B (Tab. 4), and a demonstration of the impact of scaling DPO reward models on performance in Tab. 3. We further study the limits of current reward models (Section 5.2) and prior test sets (Section 5.3).https://huggingface.co/datasets/allenai/reward-bench-results
Tab. 2 summarizes results for the top 20 models across different model sizes large, medium, and small. The large models are the only models capable of consistent high performance on the Chat Hard and Reasoning sections, with the model Starling-RM-34B (81.5) being state-of-the-art. These models are not accessible for many people to use, so we define two other categories of state-of-the-art, 7 billion parameters and 1.5 billion parameters or less. The leading medium-sized 7B models are Starling-RM-7B-alpha (74.7), zephyr-7b-alpha (73.6), and Nous-Hermes-2-Mistral-7B-DPO (73.5) given the similar scores and no formal notion of error bars on the benchmark. The final category is comprised of the small, most accessible models, where the state-of-the-art models are stablelm-2-zepyhr-1_6b (65.9) and oasst-rm-2.1-pythia-1.4b-epoch-2.5 (65.1). There are striations in performance with changes in base model size and quality, mirroring the benchmark performance of models such as OLMo, Llama 2, Mistral 7B, Yi-34B, and others.
In our evaluation there are multiple models trained either with the same or very similar fine-tuning approaches on different base models. We show the impact of scaling across different Llama 2 and Qwen 1.5 versions in Tab. 3. In general, Llama 2 shows a clear improvement with scaling across all sections of RewardBench, but Qwen 1.5 shows less monotonic improvement (and even regression on Prior Sets).
Tab. 4 compares the impact of different base models and subtle changes of fine-tuning methods via the Zephyr-class models (Tunstall et al., 2023). zephyr-7b-beta, zephyr-7b-alpha, zephyr-7b-gemma-v0.1, and tulu-2-dpo-7b are all trained with the same target method and different base models or datasets. zephyr-7b-alpha and zephyr-7b-beta differ by filtering of the UltraFeedback preference dataset only, and this is reflected in zephyr-7b-alpha’s higher score on Safety (as refusals were removed from the dataset) and lower score on Chat. tulu-2-dpo-7b shows the difference from the Mistral 7B to the Llama 2 7B base models and a different supervised fine-tuning dataset, as regressions on Chat Hard and Reasoning, but improvements on Safety. zephyr-7b-gemma-v0.1 shows the regression when switching to Gemma base model across many categories.
The per-prompt scores demonstrate the different magnitudes and distributions of rewards assigned to each reward model over the RewardBench evaluation dataset. In Fig. 2 these distributions are shown for the reward models trained as classifiers we evaluated, with results for DPO models and on prior preference test sets in Appendix A.2. Only some reward models are Gaussian in their scores, only some reward models are centered around 0 reward, and few are both. While much reward model research focuses on mitigating overoptimization, future work should identify a practical RM output distribution for downstream RL training.
2 Limits of Current Reward Models
A summary of performance is shown in Tab. 9. Current reward models can solve some subsets of RewardBench reliably, approaching 100% accuracy, but many subsets experience a combination of low ceilings on performance or high variance of performance. The subsets with low ceilings, mostly in the Chat Hard and Reasoning sections indicate areas where preference datasets and reward modeling methods can be extended to improve performance, and subsets with high variability, such as many of the Safety subsets, indicate areas where best practices can be converged upon.
Tab. 5 compares different rewards models across Chat Hard categories (full results are shown in Tab. 9). The adversarial subsets from LLMBar are crucial to understanding RMs because they show examples where two answers are written in a similar style (e.g. the same GPT-4 model version), but with slightly different subjects. The difference between asking a factual question about a related but different object or slightly changing the context of a prompt, is hard to pick up with most reward models. The Chat Hard section (and to some extent Reasoning) is the mirror of the Prior Sets section, where the hard prompts are dominated by DPO models – even those with low average performance overall, such as the Qwen Chat Models. The performance gain of DPO models can be caused by many aspects, ranging from better base models to more aligned training datasets, but closing the gap with standard reward models trained as classifiers is an important step.
The Reasoning section of RewardBench has the widest, smooth variation in performance – e.g. models populate many levels, from 35% accuracy (well below random) all the way to 90% accuracy. Though, the ceiling on reasoning models is much harder than the adversarially designed data, indicating RMs can reliably identify known bugs in reasoning or code. Full reasoning results are included in Tab. 11.
Tab. 6 (full results in Tab. 10 in Appendix) compares different reward models across different safety categories, indicating challenges on striking a balance between refusing too much or not refusing. Models, such as zephyr-7b-beta and zephyr-7b-gemma-v0.1 show how a model focused on helpfulness without a strong notion of safety will score poorly on the should-refuse subsets of the safety section, but highly on XSTest Should Respond. Other models, namely those at the top of the overall leaderboard, clearly include safety information in the training process and maintain strong performance on trick questions that could induce false refusals (XSTest Should Respond). Finally, the third option is also represented in models – those that score highly on prompts that they should refuse and poorly on those they should not, indicating a model that is likely to falsely refusal queries (for example, the Qwen chat models). These three behavior modes being represented indicates that RewardBench can be used as a quick check of the safety behavior of a candidate model, especially when trained with DPO (as it will not need further RL training like the classifier models).
Given the results showing length bias in RLHF and reward models (Singhal et al., 2023), we designed RewardBench so that the chosen responses are either a similar length or shorter than the rejected responses. For example, the AlpacaEval Length subset is designed to differentiate between other Chat subsets by having notably different models capabilities with the same average length (results in Tab. 8). In this case, the results are lower than other easy chat subsets, but 90% plus accuracy is achieved by over 10 models – far above random for most models. Though, more detailed statistical tests are needed to fully understand this, as this only tests the reward models’ abilities to discern information without the help of length as a proxy. More details on the length distributions of RewardBench are found in Appendix D.2.
3 Limitations of Prior Test Sets
Many popular models trained with RLHF use new preference datasets such as UltraFeedback (Cui et al., 2023) or Nectar (Zhu et al., 2023a), which don’t have publicly available validation sets. Given this, when training reward models, common practice is to compare model agreement with a variety of existing test sets from earlier work in RLHF.
Some models scoring strongly on the Prior Sets section of RewardBench, such as UltraRM-13b and PairRM-hf were trained on the training splits of Anthropic HH, Stanford Human Preferences (SHP), and OpenAI’s Learning to Summarize, but other top classifier models, such as the Starling models were not. Combining this with the very low average score of DPO models on these test sets indicates that substantial research is needed to understand the full limitations of these datasets. Full results are detailed in Tab. 12.
Additional data is included in the code-base, but not included in the evaluation score due to noisy results or lack of clear use instructions (e.g. could be easy for unintentional test-set contamination). In this vein, results on SafeRLHF (Dai et al., 2023) data and MT Bench labelshttps://huggingface.co/datasets/lmsys/mt_bench_human_judgments (from humans and GPT-4) are supported within the methodology, but not included in this analysis.
Discussions
Since DPO-trained LLMs are implicit reward models largely used for their generative abilities, the question of how they compare to RMs trained as classifiers is unstudied. There are currently more DPO models released to the public, partially due to DPO requiring notably fewer computational resources among other factors such as existing implementations and relevant datasets. We see that the results on RewardBench flatter the recent DPO methods, except for the Prior Sets section. For how the DPO reward is computed, see Sec. 3.
The same inference code of popular DPO training implementations can easily be used for evaluation as an RM by not propagating gradients through the models. The simplest implementations requires more GPU memory to run evaluation of DPO-trained models given the two models needed to compute the reward, but this can be avoided by computing the probabilities over the policy and base models sequentially. Though, some of the released DPO models do not clearly document which reference model is used in training (e.g. if it is a base model or a model obtained via supervised fine-tuning), which can result in unclear benchmarking.Examples include Mixtral-8x7B-Instruct-v0.1 or the Qwen chat models, which just say “trained with DPO,” yet they achieve solid performance. When a reference model is unavailable or compute is constrained, an alternative approach in such cases would be to obtain a reference free reward: , which could be normalized using different approaches. Without normalization, the loss has a length penalty by summing over probabilities of each token which are all negative numbers. We will explore the impacts of reference free inference in future work.
We also experimentedwith using the “wrong” reference model, i.e. a similar but different base model, and found that this reduced the DPO trained RM performance to similar levels as the random baseline.
There is still a lot that is unknown about the best practices of training RMs: trained with DPO they are regularized by KL distance, but the classifiers are not. Additionally, a common practice for training RMs via classification is to train for 1 epoch (Ouyang et al., 2022), while DPO models are usually trained for more than 1 epoch (Tunstall et al., 2023; Ivison et al., 2023). Other future work ideas therefore include analyzing the role of the training hyperparameters in DPO training and RM classification performance (such as Beta KL regularization on generated text, number of training epochs, etc.).
Given LLM-as-a-judge’s prevalent use for evaluation, recent works have emerged using LLMs as feedback mechanisms very similar to reward models. Some works have fine-tuned models specifically for the task of rating or choosing responses from LLMs (Jiang et al., 2023b; Kim et al., 2023; Zhu et al., 2023b). Other work has proposed generative reward modeling (Li et al., 2023a)– using a generative language model to provide scores via output tokens. While similar to the reward computation of DPO models, this mode of score calculation often involves specific prompting per-model and more computation per sample, such as explaining reasoning before or after the score. Given these differences, we decided not to include them in the RewardBench leaderboard, but they are worth exploring in future work.
Reward models inhabit an important normative role in the RLHF process being the primary artifact where human preferences or values are encoded in the final policy. The RewardBench infrastructure enables asking basic questions when studying reward models such as whose or which values are embedded as the sense of reward (Lambert et al., 2023). Initial work is studying this question for LLMs broadly, such as measuring representation (Durmus et al., 2023; Ryan et al., 2024) or moral foundations of LMs (Abdulhai et al., 2023), but this work should be extended to reward models. This can involve the study of different base models which RMs are trained from, tweaking fine-tuning techniques, if synthetic datasets amplify bias in RMs as well (Wyllie et al., 2024), and datasets.
An emerging trend in LLMs is the shift from chat systems being only a model to being a system of models, with small models used as classifiers for tasks such as safety (Mozes et al., 2023). If some LLMs or RMs are designed to be used with additional safety classifiers after the fact, evaluating them on RewardBench may not be a fair comparison. For systems such as this, each classifier for a specific task should be evaluated on the sections it controls. The most common area where this is handled is safety, where a small reward model can be used to permit or block all outputs from a larger generating model.
Conclusion
We present RewardBench, and show the variety of performance characteristics of current reward models in order to improve understanding of the RLHF process. While covering a wide variety of topics important to alignment of language models, a crucial next step is needed to correlate performance in RewardBench to downstream performance of a model trained with RLHF. We have taken a first step to understanding which values are embedded in the RLHF training and data, showing trends across many base models and preference datasets. The toolkit we have released can easily be expanded include new custom dataset to specifically audit a certain property of the RLHF process. RewardBench is one of many tools which will help us understand the science of whose and what values are embedded in our language models.
Acknowledgements
The authors would like to thank Thomas Gilbert for early discussions that helped motivate this project. Thanks to Prasann Singhal for discussing similar and complimentary concurrent work when building this project. Thanks to Hamish Ivision for helping with the math data filtering code. Thanks to Matt Latzke for help with the logo and design artifacts.
References
Appendix A Additional Results
Table 7 shows the full results for the first reward models we collected in this work. In addition, Tables 8-12 provides the performance breakdown per category.
The full distribution of accuracies for models tested on RewardBench are shown in Fig. 3 for the core dataset and in Fig. 4 for existing preference sets. The subsets created for RewardBench show substantial higher variance and range than the existing test sets used to evaluate reward models. A higher range of evaluation signal indicates that the benchmark makes it easier to differentiate between two similar models. Important subsets to RewardBench are those with maximum performance below 100%, indicating potential future work.
A.2 Model Reward Distributions
An interesting detail that is not yet easy to apply to training better RLHF models is the shape of the distribution of given reward models on the same input dataset. For all the datasets tested in RewardBench, we record the outputted scores for every prompt. The outputs of models trained with DPO are all large negative numbers given they are summations of logprobs across the generation. The outputs of reward models trained as a simple classifier should in concept be near to a unit Gaussian given desirable properties of a reward function for RL algorithms, but this is normally not the case. The distribution of the classifier models is shown for the core evaluation set in Fig. 2 and over the previous test sets in Fig. 7. The distributions for models trained with DPO are shown in Fig. 5 for classifiers and in Fig. 6 for models trained with DPO.
The custom classifiers, such as PairRM and SteamSHP are omitted because their intended use is to take two responses in at once, so a score does not apply in the same way.
Appendix B Dataset Details
Here, we detail the curation process of every subset. All subsets are either manually verified or are curated from previous evaluation datasets with manual verification. For detailed data processing notes, see Appendix E. In total there are 2958 prompts in RewardBench. All subsets in the primary dataset are single-turn instruction following tasks.
This section is designed to evaluate the basic instruction following understanding within a reward model.
Manually verified prompt-chosen-rejected trios from AlpacaEval (Li et al., 2023b) where the chosen and rejected responses come from models of different capabilities.
For the AlpacaEval Easy subset with 100 prompts, the chosen completions are from the GPT4-Turbo responses (97.70% win rate) and the rejected come from a much weaker model, Alpaca 7B (Taori et al., 2023) (26.46% win rate).
For the AlpacaEval Length subset with 95 prompts, we seek two models with similar average completion length and a large delta in evaluated performance. It is seeded from Llama 2 Chat 70B (92.66% win rate, 1790 average character length) (Touvron et al., 2023) and rejected is from Guanaco 13B (52.61% win rate, 1774 average character length) (Dettmers et al., 2023).
The AlpacaEval Hard subset contains 95 manually verified prompt-chosen-rejected trios where the chosen responses come from the Tülu 2 70B DPO responses (95.03% win rate) and the rejected come from a weaker model, Davinci003 (Ouyang et al., 2022) (50.00% win rate).
The MT Bench Easy subset is composed of 28 manually verified prompt-chosen-rejected trios from MT-Bench (Zheng et al., 2023) where chosen and rejected correspond to judgements of score 10 and 1 respectively for the same prompt.Data is available here: https://huggingface.co/spaces/lmsys/mt-bench/blob/main/data/mt_bench/model_judgment/gpt-4_single.jsonl The MT Bench Medium subset is similar, with 40 manually verified prompt-chosen-rejected trios from MT-Bench (Zheng et al., 2023) where chosen and rejected correspond to judgements of score 9 and 2 to 5 respectively for the same prompt.
For all MT-Bench subsets, the second turn data was not included due to the out-of-distribution nature for a reward model, where the data would be different across the entire conversation and not just the last turn after the prompt. Second, organizing by scoring is difficult due to scores being assigned both for the first and second responses. Further MT-Bench filtering data, such as the models included and distribution of scores, is included in Sec. E.2.
B.0.2 Chat Hard Subsets
This section is designed to challenge the instruction following abilities of a reward model with trick questions and minor factual or formatting issues.
37 manually verified prompt-chosen-rejected trios from MT-Bench (Zheng et al., 2023) where chosen and rejected correspond to judgements of score 7 to 8 and 5 to 6 respectively for the same prompt.
The 100 examples from LLMBar Natural split have preferred completions from existing instruction following benchmarks, which are manually verified in preference ranking (Zeng et al., 2023). This subset is similar to AlpacaEval and MT-Bench subsets.
Human-curated trick instruction-following questions for LLM-as-a-judge applications from LLMBar (Zeng et al., 2023) reformatted as prompt-chosen-rejected trios. Neighbor creates a rejected completion from a closely related instruction in the dataset, GPT4Inst creates a rejected by asking GPT4 for a similar instruction to the original which is then used as a generation, GPT4Out creates a rejected sample by asking GPT4 to be unhelpful when following the same prompt, and Manual is a set of specifically curated trick pairs.
The counts per subset are 134 for Neighbor, 92 for GPTInst, 47 for GPTOut, and 46 for Manual.
B.0.3 Safety Subsets
This section is designed to evaluate the propensity for reward models to prefer refusals to sensitive questions or to prefer responses to questions which could trigger a false refusal.
100 examples in each subset with prompts from GPT-3.5 and GPT-4, seeded with human-written prompts designed to elicit dangerous or offensive responses. The chosen completions are refusals from GPT-3.5, which we find to give more varied and detailed refusals than GPT-4. The rejected completions are responses that have been manually verified to contain dangerous or offensive content, sourced from Dolphin-2.0-mistral-7bhttps://huggingface.co/cognitivecomputations/dolphin-2.0-mistral-7b, an uncensored fine-tune of Mistral 7B (Jiang et al., 2023a).
136 examples from the original 900 examples in the Do Not Answer dataset (Wang et al., 2023), designed to have questions that only responsible language models will not answer, where the rejected response is an answer flagged as harmful in their analysis and the chosen is a harmless answer for the same prompt. For this subset, when there are multiple options of models to select from, we prioritize selecting a non-GPT4 model enhance dataset variety. Otherwise, the samples are randomly sampled from the suitable choices.
250 manually selected triples from XSTest (Röttger et al., 2023). The prompts from this XSTest subset were designed to test for exaggerated safety behavior by containing certain, seemingly unsafe, trigger words, while actually being safe to comply to. The chosen and rejected completions were selected from GPT-4 (Achiam et al., 2023), Llama-2-70b-chat-hf (Touvron et al., 2023) (with the original system prompt and without), Mistral-7B-Instruct-v0.1 (Jiang et al., 2023a) (with and without a guardrail system prompt), dolphin-2.0-mistral-7b and zephyr-7b-beta (Tunstall et al., 2023). In cases, when none of the model generated a refusal, we used “Sorry, but I can’t assist with that.” as the refusal text. For the should-respond subset of the queries, the rejected completion is always a refusal (or partial-refusal) and the chosen a compliance. In order to keep this test set challenging, we opted for keeping the completion source varied. Whenever no refusal was generated by either of the models, we used the following refusal template: “Sorry, but I can’t assist with that.”
154 (out of 200For 46 prompts none of the models complied and it was not feasible to get human written toxic content.) manually selected triples from XSTest (Röttger et al., 2023). For the should-refuse subset of the queries, the rejected completion is always a compliance and the chosen a refusal (or partial-refusal). The completions were selected from the same set of models as mentioned above for XSTest should-respond and we applied the same design decisions. Additionally, when no compliance was available from our set of models and it seemed feasible, we also hand-wrote some of the completions.
B.0.4 Reasoning Subsets
This section is designed to evaluate specific reasoning abilities such as code and math.
For each programming language, there are 164 prompts with buggy and functional solutions in HumanEvalPack (HEP) (Muennighoff et al., 2023). We format these with the chosen answer as the correct solution and the buggy answer as rejected.
We filter and select answers from the PRM800kPRM: process reward model. reasoning dataset (Lightman et al., 2023) to construct pairings of reference answers with incorrect, generated answers from an GPT4 fine-tune used in the paper. We use the test set from phase 2 of the data for these rollouts, filtering for examples only where the model generated an error (no doubly correct examples). The questions originate from the MATH dataset (Hendrycks et al., 2021).
Appendix C Discussion on Prior Test Sets
The goal in choosing the subsets for the Prior Sets section of the benchmark is to include results that are representative of past attempts in reward modeling and still useful to future work. Many of the datasets in this section differ from other popular preference datasets by being populated by human labels. We primarily chose to include the data for this section based on a process of elimination after evaluating many models in order to create a leader-board ranking which was fair. For example, we decided that the Safety section better represented models’ abilities. The SHP data we include is a filtered version of their subset to increase the margin between ratings, so that the data should be easier to discerne by the RMs. Full data for this section is shown in Tab. 12. The MT Bench data included in the table is interesting, but isn’t formally released as a test set, so we are worried about potential contamination (and MT-Bench is already heavily covered by the benchmark). It does, though, show interesting correlations between the agreement of human and GPT4 judgements.
Appendix D Dataset Characteristics
The following subsections will discuss our analyses of some high-level characteristics of the evaluation dataset.
Figure D.1 shows the sources of all completions in the evaluation set, whereas Figure LABEL:fig:source_completion_chosen_rejected shows the breakdown for both chosen and rejected completions. The unknown label applies to instances of LLMBar and PRM800k. For LLMBar, the authors manually filtered and modified each example to ensure their difficulty, resulting in instances that are neither fully human-generated nor fully model-generated. For PRM800k, all unknown instances are rejections because we only filtered on cases where the model generated an error.
Reward models tend to correlate reward with prompt length (Singhal et al., 2023), and so we looked into the prevalence of this bias in our preference data. For a given dataset, we measured the average prompt length (in terms of subtokens) of the chosen and rejected completions. Figure 9 shows the results.
In this section, we’ll detail our notes from the data filtering process with examples of verified and rejected prompt-chosen-rejected triples. More details are included for the AlpacaEval and MT-Bench subsets due to their more subjective nature.
The instructions used to see curating the data were as follows:
For all the categories presented below, we will manually verify all of the chosen-rejected pairs to a minimum criteria of correctness. In this process, it is better to have fewer samples than contradictory data, which reduces the signal of the benchmark. Some subsets, such as LLMBar, are filtered by the previous authors. Further filtering was conducted by multiple people following the following guidelines:
When sampling a dataset, do not skip because it is a hard choice. This will bias the subsets into being artificially easier for the reward models to understand. Rejecting due to both being wrong is common.
Follow basic principles of what makes a chatbot useful. The capabilities sets prioritize helpfulness, factuality, and honesty (similar to early work from Anthropic and InstructGPT). Harmful content could be what is requested, but I do not expect this.
When in doubt, ask for help. This is not a maximum throughput exercise. Ask on slack or email if there is a point we should discuss.
For capabilities, refusals cannot be in the chosen. For harm / safety, refusals are expected to be in the chosen.
E.2 MT Bench filtering
As discussed in the paper, our MT Bench subsets are derived by pairing higher scoring model responses with lower scoring model responses for a given prompt into chosen and rejected pairs, respectively.
Next, we manually verified all of the samples, about 10% of the completions were thrown out. We found some common trends:
Very low GPT-4 scores were often caused by gibberish / repetitive text.
Some factual verifications were needed to filter the data.
The ‘hard’ subset mostly entailed style differences, e.g. short vs. long answers, and we did not editorialize what is right as long as there was a reason.
The models used in the subsets of RewardBenchfrom MT-Bench are as follows, and of high diversity:
Models chosen: Llama-2-70b-chat, tulu-30b, guanaco-65b, vicuna-7b-v1.3, oasst-sft-7-llama-30b, Llama-2-13b-chat, gpt-4, claude-v1, mpt-30b-chat, gpt-3.5-turbo, guanaco-33b, palm-2-chat-bison-001, Llama-2-7b-chat, claude-instant-v1.
Models rejected: vicuna-7b-v1.3, wizardlm-13b, falcon-40b-instruct, rwkv-4-raven-14b, vicuna-13b-v1.3, fastchat-t5-3b, stablelm-tuned-alpha-7b, llama-13b.
Subset 2: Medium, 9s vs 2-5s (for balancing available data)
Models chosen: mpt-30b-instruct, baize-v2-13b, claude-instant-v1, wizardlm-30b, guanaco-65b, nous-hermes-13b, gpt4all-13b-snoozy, claude-v1, vicuna-33b-v1.3, mpt-7b-chat, vicuna-7b-v1.3, oasst-sft-7-llama-30b, palm-2-chat-bison-001, Llama-2-7b-chat, koala-13b, h2ogpt-oasst-open-llama-13b, vicuna-13b-v1.3, gpt-3.5-turbo, alpaca-13b.
Models rejected: mpt-30b-instruct, oasst-sft-4-pythia-12b, dolly-v2-12b, falcon-40b-instruct, gpt4all-13b-snoozy, rwkv-4-raven-14b, chatglm-6b, fastchat-t5-3b, koala-13b, alpaca-13b, stablelm-tuned-alpha-7b, llama-13b, h2ogpt-oasst-open-llama-13b.
Models chosen: baize-v2-13b, mpt-30b-instruct, rwkv-4-raven-14b, wizardlm-30b, llama-13b, oasst-sft-4-pythia-12b, tulu-30b, guanaco-65b, nous-hermes-13b, falcon-40b-instruct, gpt4all-13b-snoozy, chatglm-6b, stablelm-tuned-alpha-7b, mpt-7b-chat, mpt-30b-chat, palm-2-chat-bison-001, guanaco-33b, Llama-2-7b-chat, koala-13b, h2ogpt-oasst-open-llama-13b, Llama-2-70b-chat, gpt-3.5-turbo, alpaca-13b
Models rejected: mpt-30b-instruct, rwkv-4-raven-14b, llama-13b, oasst-sft-4-pythia-12b, guanaco-65b, falcon-40b-instruct, gpt4all-13b-snoozy, claude-v1, chatglm-6b, vicuna-33b-v1.3, stablelm-tuned-alpha-7b, mpt-7b-chat, mpt-30b-chat, palm-2-chat-bison-001, koala-13b, dolly-v2-12b, vicuna-13b-v1.3, fastchat-t5-3b, gpt-3.5-turbo, alpaca-13b
The distribution of scores in the MT Bench ratings dataset is shown in Fig. 10.
Examples from the MT-Bench Medium subset are shown in Fig. 11 (accepted) and Fig. 12 (removed).
Examples from the MT-Bench Hard subset are shown in Fig. 13 (removed for accuracy).
E.3 AlpacaEval filtering
To review, the AlpacaEval subsets are all initialized by two pairs of models (rather than scores like MT-Bench). With this in mind, filtering is still very familiar to those splits.
Some notes on errors present in the dataset prompting removal are:
Both answers, or just the chosen answer, are incorrect.
Clear formatting instructions are ignored.
Examples from AlpacEval Easy are shown in Fig. 14 and Fig. 15, which were both rejected for factual errors.
An example of a rejected sample from AlpacaEval Hard is shown in Fig. 16 and an example rejected from AlpacaEval Length is shown in Fig. 17.
E.4 Refusals data
The goal of this section was to have a refusal as the chosen and then harmful or unsafe text as the rejected option. Given a set of prompts, we sought out generating the rejected completions with various ML models. Based on evaluations in Ivison et al. (2023), we initially tried to generate harmful text with the Zephyr- model, but it proved to be inconsistent. Upon switching to the model dolphin-2.0-mistral-7b, the results were far more stable.
E.5 XSTest filtering
In order to create this dataset, we manually read the completions in the original XSTest repositoryhttps://github.com/paul-rottger/exaggerated-safety and selected one response that complied with the instruction and one that erroneously refused. For some prompts, we needed to use an additional model (again dolphin-2.0-mistral-7b) to generate prompts that complied with the prompt. After adding in dolphin-2.0-mistral-7b, there were still 46 of the original 450 prompts that did not have a suitable chosen-rejected pairing. Some of the datapoints are tricky, and we included the intent of the original authors.
For example, the text asking about violence in a fictional story is determined as one that should be responded to, which is shown in Fig. 19. On the other hand, some rather odd situations also should be answered, such as in Fig. 18.