Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Haoxiang Wang, Wei Xiong, Tengyang Xie, Han Zhao, Tong Zhang
Introduction
In this paper, we explore the role of reward models (RMs) within the framework of Reinforcement Learning from Human Feedback (RLHF). RMs play a crucial role in aligning large language models (LLMs) as they provide a scalable way to integrate human preferences into the models’ training process, guiding the optimization of their policies. To be more specific and provide more context, we first review the most standard and popular RLHF frameworks and the role of RMs in this framework. Arguably the dominant RLHF approach is a deep reinforcement learning (DRL)-based framework, as developed in key studies [Christiano et al., 2017; Ouyang et al., 2022; Bai et al., 2022]. This framework operates in three stages: 1) Preference data collection; 2) Reward modeling based on the Bradley-Terry model [Bradley and Terry, 1952]; 3) Policy optimization using Proximal Policy Optimization (PPO) [Schulman et al., 2017] and the reward model constructed in stage 2. This framework has achieved tremendous success in the post-training of ChatGPT [Ouyang et al., 2022] and Claude [Bai et al., 2022]. These ideas also extend to other approaches, such as rejection sampling fine-tuning [Dong et al., 2023; Gulcehre et al., 2023] and iterative direct preference learning [Xiong et al., 2023; Guo et al., 2024; Xie et al., 2024]. In these approaches, the intermediate policy is typically iteratively deployed to collect new responses, uses the reward model to label the responses, and fine-tunes the model on the newly collected preference data. In all of these RLHF frameworks, the capacity of the reward model is crucial as it directly affects the quality of the aligned LLMs.
The most popular reward modeling approach is based on the maximum likelihood estimation (MLE) of the Bradley-Terry (BT) model [Bradley and Terry, 1952]. Despite its widespread use, the BT model is rather limited in the capacity of capturing the complicated human preference [Munos et al., 2023; Swamy et al., 2024; Ye et al., 2024]. In addition to the capacity issue, common RMs, like the BT model, are typically black-box models that output scores or preferences without providing human-interpretable explanations, making it subject to the widely observed phenomenon of reward hacking [Skalse et al., 2022; Singhal et al., 2023; Chen et al., 2024], where the aligned LLMs generate high-reward responses (rated by the RM) that do not align with actual human preferences [Gao et al., 2023; Lin et al., 2023; Coste et al., 2023]. A notable example of this is the verbosity bias, where aligned LLMs produce longer-than-necessary responses because the RM favors length, regardless of quality [Singhal et al., 2023; Wang et al., 2024a; Chen et al., 2024].
In this work, we aim to enhance reward models by making them more interpretable [Molnar, 2020] and steerable [Wong et al., 2021]. Using the aforementioned verbosity bias as an example, suppose the RM’s output is decomposable, meaning that it assigns a high score to a response due to two factors: 40% for its helpfulness and 60% for its length. In this case, we can see that the RM may suffer from the verbosity bias. Furthermore, if the RM is steerable, we could adjust its decision-making process to base its scoring 100% on helpfulness. This would be regardless of response length, thus mitigating the verbosity bias. Enhancing the interpretability of RMs also allows humans to verify whether RMs have similar internal decision processes to humans when acting as proxies for human preferences. We believe that this human-AI interaction process could ensure that RMs are consistent with human values and preferences, making RM-aligned LLMs more reliable and robust.
At a high level, we propose a two-stage approach that first trains a multi-objective RM and then learns a gating layer that scalarizes reward objectives in a mixture-of-experts way. We then empirically validate its effectiveness by training such an RM with Llama-3 8B [Meta, 2024], and obtain state-of-the-art performance on RewardBench, a benchmark to evaluate RMs.
Related Works
The PPO-based RLHF framework is first popularized in Christiano et al. and further developed by Bai et al. ; Ouyang et al. to make ChatGPT and Claude, which leverages a reward model to provide feedback during the RLHF process. However, getting the PPO work is challenging in the context of LLMs [Choshen et al., 2019; Engstrom et al., 2020]. Thus, much efforts have been made in proposing alternative approaches to the PPO, such as the REINFORCE algorithm variants [Li et al., 2023; Shao et al., 2024]. Another popular approach is the reward-ranked fine-tuning algorithm (RAFT) [Dong et al., 2023; Gulcehre et al., 2023] that was used in LLaMA2 [Touvron et al., 2023], Llama-3 [Meta, 2024], Qwen2 [qwe, 2024] and Apple Intelligence. To implement rejection sampling, we typically sample responses per prompt and use a reward model to rank them according to some criteria. Then, we fine-tune the model on the high-rank responses (e.g., the one with the highest reward value). This algorithm is a strong baseline, especially in reasoning tasks [Aksitov et al., 2023; Havrilla et al., 2024]. All approaches mentioned above leverage external reward models to provide supervision signals during the RLHF process.
There is also a line of works studying direct preference learning algorithms [Zhao et al., 2023; Rafailov et al., 2023; Azar et al., 2023; Tang et al., 2024], which bypasses traditional reward modeling to learn directly from preference datasets in a supervised manner (hence the name direct preference learning). Direct Preference Optimization (DPO) is the most representative one. However, the original DPO is an offline algorithm without further exploration of the environments. The subsequent studies demonstrate that the online iterative variants surpass the original DPO with large margins [Xiong et al., 2023; Liu et al., 2023; Xu et al., 2023; Rosset et al., 2024; Guo et al., 2024; Xie et al., 2024; Zhang et al., 2024; Dong et al., 2024]. Specifically, we can iteratively deploy the intermediate policy to collect new responses and use the external reward model to label them, and further fine-tune the model on the newly collected preference data using the DPO objective.
To summarize, all the existing popular RLHF algorithms require an external reward model to provide preference signals to achieve their best performance.
2 Reward modeling in RLHF
Traditionally, reward models in RLHF have utilized the Bradley-Terry (BT) model for preference estimation [Bradley and Terry, 1952; Ouyang et al., 2022; Bai et al., 2022; Wang et al., 2023b; Rafailov et al., 2023]. Despite its widespread use, the BT model’s inability to handle complex, in-transitive preferences has been highlighted in recent studies [Munos et al., 2023; Swamy et al., 2024; Ye et al., 2024]. It is also argued that the DPO-aligned model can serve as a reward function to provide token-wise rewards [Rafailov et al., 2024; Zhong et al., 2024], which are still confined to the BT model. There are also works dropping the BT assumption and directly modeling the probability of response one being preferred over another one [Jiang et al., 2023; Zhao et al., 2023; Liu et al., 2023; Dong et al., 2024]. These models are referred to as the pairwise preference model, as they take two responses as the input. Another line of work explores multi-objective reward models that attempt to capture the complicated human preferences more effectively [Touvron et al., 2023; Wang et al., 2023a, 2024a]. However, the integration of these multi-dimensional signals typically relies on naive methods such as linear combinations, indicating a need for more sophisticated techniques.
Methodology
Most existing reward models for LLM alignment are trained with Bradley-Terry loss on pairwise data with annotated preferences [Bai et al., 2022; Touvron et al., 2023; Ouyang et al., 2022], using the same approach as InstructGPT [Ouyang et al., 2022]. The pairwise preference annotations are essentially binary labels, e.g., , indicating which response is preferred by the annotator. We call them relative ratings here. However, in some recent high-quality datasets, the relative ratings are converted from absolute ratings. For instance, UltraFeedback [Cui et al., 2023] is curated with 5-objective absolute ratings: Overall Score, Instruction Following, Truthfulness, Honesty, and Helpfulness (each objective has 5 distinct ratings based on pre-defined rubrics). The dataset is further binarized into pairwise comparisons, using the Overall Score, or the average score of the remaining 4 objectives, for training reward models or DPO. The original ratings are fine-grained, as each objective has continuous integer rating scores (e.g., 1, 2, 3, 4, 5). However, the binarization process discards some fine-grained information. For example, a pair of examples with scores 1:5 is labeled in the same way as another pair with scores 2:3. It is not justified that discarding the fine-grained preference information is beneficial. Hence, we would like to include all fine-grained information for reward modeling.
2 Mixture-of-Experts Scalarization of Reward Objectives
An ArmoRM can predict multi-objective rewards for each response. However, the multi-dimensional outputs need to be reduced to a scalar for ranking or pairwise comparisons of test examples. A straightforward approach is to take a linear combination of multiple objectives [Hu et al., 2024] as in the literature of multitask learning. However, using fixed combination coefficients is too rigid for complex application scenarios. For instance, for prompts that could easily trigger unsafe responses, the safety objective should be assigned a large coefficient, as we wish the reward model to rank unsafe responses lower than safe ones. For prompts for math problem assistance, the safety objective becomes less relevant, and the helpfulness-related objectives should be the primary focus.
The gating layer can simply be a shallow MLP (i.e., fully-connected network) that takes the prompt feature and outputs a -dimensional vector, followed by a softmax function to ensure the elements of the output vector are non-negative and summing up to 1.
However, most reward objectives are highly correlated with verbosity, which indicates a strong verbosity bias [Saito et al., 2023]. Using non-negative gating coefficients would make the final output inherit the bias. To resolve the issue, we adjust each reward objective, , with a penalty using the verbosity reward objective,
where the penalty coefficient is chosen such that for a proper correction metric (e.g., Pearson or Spearman correlation coefficient) and a reference data distribution ,
Finally, we multiply the gating coefficients to the multi-objective rewards, to obtain a scalar score for the response given prompt
Experiment
We use the Llama-3 8B [Meta, 2024] architecture and initialize the model backbone with parameters from a Bradley-Terry RM of Llama-3 8B trained by Dong et al. . We append a linear layer to the backbone, and train it with regression loss while keeping the backbone frozen. The training involves 19 objectives (including helpfulness, correctness, verbosity, etc.) from 8 datasets, with details presented in Appendix A.
Implementation of MoE
Software
Our training code is built with PyTorch [Paszke et al., 2019], HuggingFace’s Transformers [Wolf et al., 2019] and Scikit-learn [Pedregosa et al., 2011].
Hardware
Training ArmoRM (the multi-objective reward modeling stage) only involves training the last linear layer (i.e., linear probing), so we save features extracted from the backbone locally and then conduct linear probing with Scikit-learn’s linear regression solver on a CPU. For the MoE stage, we also save features locally, and then train the gating layer on a single NVIDIA A6000 GPU.
Hyperparameters
The gating layer is trained using the AdamW optimizer [Loshchilov and Hutter, 2019] with a learning rate of 0.001 for 10,000 steps with a batch size of 1024. We also apply a cosine decay learning rate scheduler.
Evaluation Benchmark
RewardBench [Lambert et al., 2024] is the first benchmark constructed to evaluate reward models for language modeling. It consists of a diverse set of tasks designed to assess the performance of reward models for LLM alignment, including four primary categories (Chat, Chat Hard, Safety, Reasoning) and a category of prior sets. Each category consists of multiple datasets with pairwise preference data, where each pair includes a chosen and a rejected text response. The overall score is computed as a weighted average over the five categories, where the four primary categories have weights 1.0 and the prior-sets category has weight 0.5.
Evaluation Results
Table 1 compares the performance of our approach (ArmoRM + MoE) against other reward models. Several key observations can be made from these results:
Our model significantly outperforms the Llama-3 8B Bradley-Terry RM, which provides the LLM backbone for our model. This demonstrates the effectiveness of our ArmoRM design and the MoE gating mechanism in improving the performance of reward models.
Our model also outperforms the LLM-as-a-judge approach [Zheng et al., 2023] with GPT-4 judges by a considerable margin, indicating that our model could be used as a cheaper replacement for GPT-4 in many annotation jobs.
Our model of 8B parameters has performance nearly on par with the Nemotron-4 340B RM Wang et al. [2024b], a giant reward model of 340B parameters. This highlights the power and potential of our reward modeling approach.
Conclusion
In this work, we addressed the critical issue of interpretability in reward models for RLHF in the context of aligning LLMs with human preferences. We proposed a novel two-stage approach, consisting of an ArmoRM and a MoE strategy with a gating network. Our ArmoRM, trained with Llama-3 8B, achieved state-of-the-art performance on RewardBench, demonstrating the effectiveness of our reward modeling approach.
References
Appendix A Experimental Details
The model we use and fine-tune follows the Meta Llama3 license. All the datasets we use are open-sourced and can be used for research purposes (some could be used for commercial purposes, such as HelpSteer [Wang et al., 2023a]).
Personally Identifying Info or Offensive Content
For all datasets used in this work, according to their data curation process descriptions, they do not contain any information that names or uniquely identifies individual people, except for some examples that contain celebrity names. However, BeaverTails [Ji et al., 2023], PKU-RLHF [Ji et al., 2023], and HH-RLHF [Bai et al., 2022, Ganguli et al., 2022] contain offensive content, which is deliberately selected to build human preference datasets that aim to teach LLMs which responses are safe to generate.
Multi-Objective Training Datasets
In the stage of multi-objective reward modeling, we use training datasets with corresponding reward objectives detailed below.
HelpSteer [Wang et al., 2023a] (35k data):
helpsteer-verbosity (This is the verbosity objective we use in Eq. (2) and (3))
UltraFeedback [Cui et al., 2023] (240k data):
BeaverTails-30k [Ji et al., 2023] (30k data):
CodeUltraFeedback [Weyssow et al., 2024] (50k data):
Prometheus [Kim et al., 2024a] (200k data):
Argilla-Capybarahttps://hf.co/datasets/argilla/Capybara-Preferences-Filtered [Daniele and Suphavadeeprasit, 2023] (15k data):
Argilla-OpenOrcahttps://hf.co/datasets/argilla/distilabel-intel-orca-dpo-pairs (13k data):
Argilla-Math-Preferencehttps://hf.co/datasets/argilla/distilabel-math-preference-dpo (2.4k data): This dataset shares the objective ultrafeedback-instruction-following with UltraFeedback
Multi-Objective Data Pre-processing
When merging multiple datasets with absolute ratings (e.g., UltraFeedback and HelpSteer), we observe some issues with the data. Here, we present the issues and our approach to tackle them:
Different Rating Scales: Different datasets may have different scales for the ratings. For instance, HelpSteer has a rating scale of 0-4, while UltraFeedback’s is 1-10. We linearly transform all ratings to make them between 0 and 1. For BeaverTails with True/False ratings (indicating safe or unsafe), we treat True as 1 and False as 0.
Similar Objectives: There are some very similar objectives from different datasets. For example, the Helpfulness objective appears in both HelpSteer and UltraFeedback, and the Correctness objective of HelpSteer is quite similar to the Truthfulness of UltraFeedback. After carefully examining the datasets, we decided to treat similar objectives as separate objectives, as they are rated by different judges following different rubrics. For instance, data from HelpSteer are rated by 200 U.S.-based human annotators following customized rubrics, and UltraFeedback data are labeled with GPT-4 following another set of rubrics.
Missing Labels of the Merged Dataset: When merging multiple datasets, each example of the merged dataset only has a subset of ratings; for example, each example from HelpSteer only has 5 ratings originating from the HelpSteer dataset, and it does not have ratings for other objectives (e.g., the objectives from UltraFeedback or BeaverTails). Hence, when optimizing the regression loss, we simply ignore the missing rating dimensions of each example and only compute the loss on the remaining dimensions.
Training Data of MoE
In the stage of the gating layer, we use the following preference datasets:
HelpSteer [Wang et al., 2023a] (37k pairs)
UltraFeedback [Cui et al., 2023] (340k pairs)
SHP [Ethayarajh et al., 2022] (93k pairs)
HH-RLHF [Bai et al., 2022, Ganguli et al., 2022] (157k pairs)
CodeUltraFeedback [Weyssow et al., 2024] (50k pairs)
PRM-Phase-2 [Lightman et al., 2023] (80k pairs)
Prometheus2-Preference-Collection [Kim et al., 2024b] (200k pairs)
Preference Data Pre-processing
For datasets that are not binarized into response pairs (e.g., HelpSteer, UltraFeedback, SHP), we take the binarized versions pre-processed in Dong et al. .