The Trickle-down Impact of Reward (In-)consistency on RLHF

Lingfeng Shen, Sihao Chen, Linfeng Song, Lifeng Jin, Baolin Peng, Haitao Mi, Daniel Khashabi, Dong Yu

Introduction

Recently, reinforcement learning from human feedback (RLHF) has emerged as a popular technique to optimize and align a language model with human preferences (Ouyang et al., 2022). RLHF provides a natural solution for optimizing non-differentiable, scalar objectives for language models, and has been the centerpiece of recent state-of-the-art large language models (LLMs) (Lu et al., 2022; Hejna III & Sadigh, 2023; Go et al., 2023; Korbak et al., 2023; OpenAI, 2023).

In RLHF, a reward model (RM) generates scalar rewards for model-generated outputs as supervision during reinforcement learning. RMs are typically calibrated to proxy human preferences/rankings of responses, in the context of an input instruction. Since policy gradient methods optimize based on this reward function, the reward function inevitably dictates the behavior of the resultant chatbot. As such, the properties of RMs and their impact on RLHF models have become points of interest for the research community (Gao et al., 2022; Zhu et al., 2023; Dong et al., 2023).

In this work, we study the phenomenon of reward inconsistency in RMs, i.e., current RMs trained with the standard ranking objective on human preference data (section 2) often fail to distinguish between more vs. less favorable responses with respect to real-world instructions. We observe that reward inconsistency has a trickle-down effect on the RLHF process — the more inconsistent the RM is, the more likely the resulting chatbot is to generate inaccurate or less useful responses.

To illustrate and quantify the degree of reward inconsistency in RMs, we propose Contrast Instructions, a simple, intuitive benchmarking strategy that can be used with any instruction-tuning or human preference dataset in a fully automated manner. Contrast Instructions involves pairs of similar instructions with different responses. As an example, consider two similar-looking prompts (I^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}} and I^{{\color[rgb]{1,.5,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,.5,0}B}}) one about “RAM” and “ROM” (Figure 1; left). Despite the similarity of these prompts, they warrant different responses. Our consistency metrics measure whether a reward model can appropriately identify the correspondence between prompts and responses. Specifically, given a pair of such (instruction, response) examples, we consider an RM consistent if it assigns a higher score to the corresponding (instruction, response) compared to the other combinations, i.e. swapping instructions or responses between the two examples, as we show with 1 in Figure 1.

We construct Contrast InstructionsThe four Contrast Instructions benchmark datasets and our code are available at https://github.com/shadowkiller33/Contrast-Instruction with four popular open-source human preference or instruction tuning datasets. Surprisingly, we observe close to random-chance performance when we evaluate standard RMs trained with ranking objectives (e.g. LLaMa-7B) on Contrast Instructions, while humans are able to rank the responses correctly in ≈80%\approx 80\% of the cases. The performance gap indicates the inherent reward inconsistency from standard RM training and inference.

To enhance the consistency of standard RM, we introduce two techniques (section 5) : ConvexDA and RewardFusion ( 2 in Figure 1). The two methods can be incorporated during RM training and inference, respectively, without incurring extra computational cost during training. While our experimental results indicate some improvements, the gap with respect to human performance remains large. Interestingly, our analysis (section 6) reveals that using a more consistent RM during RLHF training leads to more useful responses from the downstream RLHF model ( 3 in Figure 1). Our findings suggest the value of reward (in-)consistency as an intrinsic evaluation metric for preference-based RMs, potentially enabling easier access for future research on reward modeling and RLHF.

The main contributions of this paper are:

We introduce Contrast Instructions, an intuitive yet scalable benchmarking strategy for evaluating RM consistency. We observe a wide performance gap between standard, preference-based RM vs. human judgments, which suggests sizable room for improvements in standard reward modeling and evaluation strategy.

We show that reward consistency can be enhanced without extra training costs. We demonstrate this with two techniques ConvexDA and RewardFusion, which can be applied during RM training and inference stages, respectively.

We empirically show that training RLHF models with more consistent RMs would result in the RLHF model generating more useful responses. We provide thorough analysis and examples, which help us understand the advantage of a more consistent RM over a less inconsistent one.

Preliminaries: Reward Modeling for RLHF

Following the conventional setup (Ziegler et al., 2019), we define RM Rθ\mathcal{R}_{\theta} as a scalar function that allows us to train a generative model during RLHF. To supervise RM, we are given a dataset of human preferences Dh\mathcal{D}_{h}. Each instance in this dataset (Ii,ri+,ri−){(I_{i},r_{i}^{+},r_{i}^{-})} is comprised of an instruction prompt IiI_{i}, a pair of responses ri+,ri−r_{i}^{+},r_{i}^{-} where ri+r_{i}^{+} is preferred over ri−r_{i}^{-} by humans. On this labeled data, Rθ\mathcal{R}_{\theta} is trained to assign a higher scalar reward to human-preferred ri+r_{i}^{+} over non-preferred ri−r_{i}^{-} in the context of IiI_{i} This can be achieved by minimizing the ranking loss L\mathcal{L}, where σ\sigma is the sigmoid function and Ii∘ri+{I}_{i}\circ r_{i}^{+} is the concatenation of IiI_{i} and ri+r_{i}^{+}.

Reinforcement Learning.

Contrast Instructions: measuring reward (in-)consistency

Conventionally in RLHF, reward models are trained explicitly to distinguish the more vs. less favorable responses in the context of an instruction (Eq. 1). One would expect an RM to consistently predict higher reward scores toward the more favorable instruction-response pairings. For instance, as the example in Figure 1 shows, an ideal reward should assign a higher score to rAr_{A} appearing in response to IAI_{A}, than IBI_{B}. In practice, however, RMs usually suffer from over-optimization towards the training distribution, leading to inconsistencies between RM predictions vs. human preferences at inference time (Gao et al., 2022).

To illustrate and measure reward inconsistency in RMs, we introduce Contrast Instructions. A Contrast Instructions benchmark takes the form of \mathcal{D}=\{(I_{i}^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}},I_{i}^{{\color[rgb]{1,.5,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,.5,0}B}},r_{i}^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}},r_{i}^{{\color[rgb]{1,.5,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,.5,0}B}})\}_{i=1}^{N} (see Fig. 1). Each test instance consists of instruction-response pairs (I^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}},r^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}}) and (I^{{\color[rgb]{1,.5,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,.5,0}B}},r^{{\color[rgb]{1,.5,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,.5,0}B}}). To make the benchmark meaningfully challenging, we sample I^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}} and I^{{\color[rgb]{1,.5,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,.5,0}B}} such that the two are lexically similar instructions with different semantics (details later in section 3.1). r^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}} and r^{{\color[rgb]{1,.5,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,.5,0}B}} are the human-preferred responses to I^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}} and I^{{\color[rgb]{1,.5,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,.5,0}B}}, respectively. Conceptually, a consistent RM should be able to identify pairs of corresponding instruction-response. Concretely, this means assuming higher score to (I^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}},r^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}}) than (I^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}},r^{{\color[rgb]{1,.5,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,.5,0}B}}) or (I^{{\color[rgb]{1,.5,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,.5,0}B}},r^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}}). There are two ways one can quantify such notions of reward consistency:

Response Consistency (Cres\mathcal{C}_{\text{res}}): Given one of the instructions I^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}}, we measure if a RM (Rθ\mathcal{R}_{\theta}) can identify the corresponding response r^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}} by assigning higher rewards to r^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}} over r^{{\color[rgb]{1,.5,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,.5,0}B}}.

Instruction Consistency (Cins\mathcal{C}_{\text{ins}}): Similarly, given one of the responses r^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}}, we measure if a RM (Rθ\mathcal{R}_{\theta}) can assign higher rewards to its corresponding instruction I^{{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}A}} over the distractor I^{{\color[rgb]{1,.5,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,.5,0}B}}.

It is worth noting that with the standard learning objective (Eq. 1), RMs are explicitly trained to rank responses but not instructions, and so Cres\mathcal{C}_{\text{res}} conceptually resembles the RM learning objective, while Cins\mathcal{C}_{\text{ins}} does not. Therefore, we expect the standard RM to perform better under the Cres\mathcal{C}_{\text{res}} metric compared to Cins\mathcal{C}_{\text{ins}}, which we will further discuss in section 4.

Contrast Instructions can be automatically constructed from datasets that contain human preferences. Inspired by Gardner et al. (2020), we formulate examples in Contrast Instructions such that the pair of instructions IA,IBI^{A},I^{B} are lexically similar, yet their ground truth responses are different. We expect a consistent RM to recognize the nuanced semantic difference between the instructions, and recognize the corresponding answer to an instruction.

We adopt four open-source human preference datasets of various NLP tasks: StackExchange for question answering (Askell et al., 2021), WMT for machine translation (Ma et al., 2019), RealSumm for text summarization (Bhandari et al., 2020), and Twitter for paraphrase generation (Shen et al., 2022c). Each dataset features examples of an instruction comprised of both a task prompt and a task-specific input, and the corresponding responses ranked by human preference. Within each dataset, we sample pairs of similar instructions with sentence embedding model SimCSE (Gao et al., 2021). To ensure the instruction pairs are similar but not semantically equivalent, we keep only instruction pairs with cosine similarity within [0.75,0.9][0.75,0.9]. We show the statistics and examples of each resulting Contrast Instructions dataset in Table 1

Inconsistency of existing reward models

With Contrast Instructions, we evaluate the consistency of RMs trained with the standard ranking objective (Eq. 1).

We initialize RMs from LLaMa-7B checkpoint (Touvron et al., 2023) and finetune on each of the four human preference datasets used to create our Contrast Instructions benchmark. We also finetune a multi-task version (Mishra et al., 2022) on the mixture of four human preference datasets. We use the following configurations for RM training for both single-task and multi-task settings. Due to resource constraints, we adopt the Low-Rank Adaptor (LoRA) (Hu et al., 2021) for training. We use the AdamW optimizer and set a learning rate of 2e-5. For multi-task training, we combine the training set from selected benchmarks and train using LoRA with a learning rate of 3e-5. Finally, we report human performance resulting from the majority vote of three human annotators (the authors) on 100 randomly selected data points.

Inconsistency of RMs on Contrast Instructions.

Table 2 summarizes the results of the finetuned RMs averaged on the four Contrast Instructions benchmarks. We observe that with a relatively large 7B parameter RM, it still performs close to random guessing in terms of both responses (Cres\mathcal{C}_{\text{res}}) and instruction consistency (Cins\mathcal{C}_{\text{ins}}). This trend can be seen on per-dataset as well, in Table 3. In our setting, multi-task training does not seem to help (no cross-task transfer), possibly due to the limited commonalities among these tasks. The reward models are slightly better in terms of Cres\mathcal{C}_{\text{res}} compared to Cins\mathcal{C}_{\text{ins}}, which fits our expectation – the RM learning objective (Eq.1) is more similar to Cres\mathcal{C}_{\text{res}} than it is to Cins\mathcal{C}_{\text{ins}}. The estimated human performance shows a wide gap compared to the RM performance on Contrast Instructions, suggesting reward inconsistency can be attributed to the standard practice of reward modeling.

Enhancing reward model consistency

Having identified in section 4 the inconsistencies of reward models, we now present methods to address these issues. Specifically, we propose two solutions: one to be applied during training (section 5.1) and another during inference (section 5.2), which does not impose additional computing costs. It’s important to note that these techniques are designed to be agnostic to Contrast Instructions’s format and setup, thereby minimizing the negative impact of Goodhart’s law (Manheim & Garrabrant, 2018).

To mitigate the effect of over-optimization during RM training, we design ConvexDA, a lightweight, efficient data augmentation technique, that neither modifies the RM learning objective nor increases the overall training cost. The high-level idea is to create various perturbations of the original input through data augmentation and select the most representative one to replace the original input example during training.

Given a human preference example (I,r+,r−)(I,r^{+},r^{-}), we use an off-the-shelf textual data augmentation tool (Ma, 2019) to substitute words in the responses with synonyms according to WordNet (Miller, 1995) or PPDB (Ganitkevitch et al., 2013). For each example, we generate N=5N=5 augmented versions {I,rj+,rj−}j=1N\{I,r_{j}^{+},r_{j}^{-}\}_{j=1}^{N}. As using all NN data points for training would incur extra cost, we draw inspiration from previous studies (Chen et al., 2010; Sener & Savarese, 2018; Rajput et al., 2019; Agarwal et al., 2020) and select one data point among five that serve as a vertex of a convex hull in the embedding space. This geometric approach ensures that we focus on the most critical points within the set of augmented data points. The strategy keeps the training efficiency on par with the standard RM training while offering the advantages of data augmentation.

We use the SimCSE model (Gao et al., 2021) to embed each original and augmented response to a 768-dimensional vector. In principle, to construct a convex hull in a kk-dimensional embedding space, we need at least kk+11 data points. For such reason, we apply Principal Component Analysis (PCA) to reduce the dimensional of the embedding to 2, We identify and select the example that act as the corner points of a low-dimensional convex hull, and replace the original input with it during RM training.

2 Consistency-Inducing Inference via RewardFusion

During RM inference, we introduce RewardFusion. The method takes inspiration from (Zhao & Cho, 2019), where we first identify sentences similar to the response from the training corpus, and then use the weighted average reward score across the target response and the retrieved training samples as a better estimate for the reward score.

At the inference time, given a pair of instructions and response I∘rI\circ r, we again use the SimCSE model to retrieve a set of similar training examples {I∘r∗}\{I\circ r^{*}\} that have cosine similarity to I∘rI\circ r over a threshold δ\delta from the training corpus. Then, we take the weighted average of the reward scores of the original plus retrieved training examples X=(I∘r)∪{I∘r∗}X=(I\circ r)\cup\{I\circ r^{*}\}:

The threshold δ=0.95\delta=0.95 is selected based on averaged performance on development sets. The embedding retrieval process is instantiated with Faiss (Johnson et al., 2019).

3 Experiments and Results

We conduct experiments to evaluate the effectiveness of both techniques in enhancing the consistency of RM on the four Contrast Instructions benchmarks. We follow the same experimental settings detailed in section 4. We start with single-task RMs trained on each human preference dataset and apply ConvexDA and RewardFusion.

We observe the following based on the results in Table 3. (1) Both ConvexDA and RewardFusion effectively enhance the consistency of the reward model. (2) The combination of these two techniques works best, highlighting their complementarity in RM training and inference for enhancing consistency. (3) Despite the improvements brought by the two techniques, the overall performance on the Contrast Instructions benchmark remains limited. RMs with enhanced consistency from the two techniques still fall behind human performance by a large margin, demonstrating the challenges of addressing RM’s consistency. We defer further details to Appendix F.

Trickle-Down Effect of Reward Inconsistentency on RLHF

Next, we explore the merits of having a more consistent RM in downstream RLHF training by comparing RLHF-trained language models with a standard (section 4) vs. more consistent RM (section 5).

We follow the overall experimental setup outlined in StackLLaMa (Beeching et al., 2023). For the experiments, we use LLaMa-7B to initialize the supervised finetuning (SFT), reward, and policy models. We train the models on StackExchange, which is segmented into SFT, RM, and RL datasets. More implementation details can be found Appendix C. We train two RLHF models, one with the standard RM, and the other with RM finetuned on examples with Contrast Instructions format sampled from StackExchange training split. We denote this as the Contrast Ft. RM. We conduct a human evaluation on 500 questions randomly sampled from the test split of StackExchange.

Human evaluation.

To assess the general quality of the responses, we conduct two evaluations on responses generated by the two RLHF models. We first ask human raters to assess the individual acceptability of each model’s response. This process involves a three-way judgment {Accept, Reject, Unsure}. A response is deemed acceptable only if it adequately addresses the question in the prompt; exhibits no significant error; and contains no redundant information. In addition, we ask the human rater to annotate the pairwise preference between the two responses. We ask human raters to compare the outputs of two models and determine which model’s output is preferred. We also provide raters the option to indicate a tie between the responses. The human evaluation results are shown in Figure 2. We observe that with the more consistent RM, the downstream RLHF model demonstrates higher generation quality.

Automatic evaluation.

We use MT-benchhttps://huggingface.co/spaces/lmsys/chatbot-arena-leaderboard (Zheng et al., 2023) for automatic evaluation. MT-Bench is a benchmark featuring challenging one-turn or multi-turn reasoning and instruction-following examples, where the model responses are graded by GPT-4 on a scale of 1 (worst) to 10 (best). Overall, the results are shown in Table 4. We observe that more consistent RMs lead to RLHF models generating more preferable responses under both settings.

Reward consistency ⇒⇒\Rightarrow More useful responses.

To understand where the improvement lies, we follow criteria from Malaviya et al. (2023) and ask human raters to assess the relevance, usefulness and factuality of RLHF model responses. Relevance indicates whether the response is topically relevant to the instruction. Usefulness indicates whether a response serves as a useful and direct answer to the instruction. Factuality indicates the factuality of responses based on the raters’ judgment, for which minimal browsing on the internet is allowed if needed. The results are shown in Table 5. We observe a statistically significant improvement in usefulness (p<0.01p<0.01 with paired t-test). Otherwise, we see minimal impact on the relevance or factuality of responses. We show a pair of example responses in Table 6, and more examples can be found in Appendix G. Overall, we observe that the usefulness of the responses from the RLHF model with standard RM tends to decrease further into the response, while the issue is mitigated with a more consistent RM.

Discussion

A reader may wonder whether the evaluation via the original datasets (task prompts and responses preferred by humans) is sufficient to reflect RMs’ level of consistency. (Alternatively, is Contrast Instructions really necessary?)

To illustrate this, we evaluate RMs’ average performance on the test set of the four human preference datasets (section 3). Table 7 shows the results. We observe that despite the improvements in terms Cres\mathcal{C}_{\text{res}} and Cins\mathcal{C}_{\text{ins}} of our ConvexDA + RewardFusion approach, the level of improvements are not reflected in the original RM evaluation metric (RMEval). To further validate this observation, we evaluate the Contrast Ft. RMs, i.e. RMs finetuned on training examples with Contrast Instructions format sampled from each of the four human preference datasets (section 6). We observe a similar pattern that despite the large (but unfair) improvements on Cres\mathcal{C}_{\text{res}} and Cins\mathcal{C}_{\text{ins}}, the performance on RMEval remains stale. These findings suggest that beyond the issue of RM over-optimization, we potentially need to rethink the current standard setup for preference-based reward modeling, which might be the inherent cause of reward inconsistency.

2 A closer look to why (or where) reward models are inconsistent

The underlying motivation behind RM consistency closely resembles model calibration (Guo et al., 2017), i.e. in the ideal case, we would expect the reward scores from an RM to be perfectly calibrated and correlated with human preferences. In such a sense, reward consistency serves as a good indicator and proxy measure for the correlation between the reward score from RMs vs. human judgments. In Figure 3, we show the reward score correlation on two human preference datasets, Twitter and RealSumm, where multiple candidate responses with their corresponding human-rated quality (scaled between 0-1) are provided for each instruction. We compare the reward score correlation from the standard reward model (i.e., left two plots in Fig 3) vs. from the RM enhanced with ConvexDA and RewardFusion (right two plots). Generally, we observe that the more consistent RM yields a higher reward score correlation with humans. While standard RMs can achieve a relatively good correlation with human preferences, the reward score exhibits higher variance, especially when it comes to pairs of responses that are closer in terms of human score. This echoes our observations with respect to standard RM learning objectives and evaluations. From RMs’ perspective, correctly distinguishing between a clearly good vs. bad response is easy. Optimizing or evaluating against such would not indicate RMs’ alignment with human preference. Including such easy pairs for evaluation would likely lead to overestimation of RMs performance, which further motivates Contrast Instructions as a benchmarking strategy for RMs, as well as a potential training strategy, as we demonstrate in section 6.

3 Consistency check beyond Contrast Instructions

Contrast Instructions provides an automatic, efficient, and intuitive evaluation framework for assessing the consistency of preference-based reward modeling. Nonetheless, it does not necessarily encompass all possible phenomena with respect to RM consistency or robustness. For such reason, we conduct preliminary analysis in adversarial and backdoor attacks (Chen et al., 2017; Shen et al., 2023) for RM. The details of the experiments are included in Appendix D and Appendix E. We observe that overall RMs suffer from a high attack success rate and exhibit vulnerability to both adversarial and backdoor attacks. The findings suggest the potential implication of reward inconsistency on RM and subsequently RLHF safety.

4 Broader Impact and Limitations

Despite the widespread interest in RLHF within the research community, our understanding so far on “what type of reward modeling would most benefit RLHF” remains fairly limited. We argue that this can in part be attributed to (1) the lack of a proper intrinsic evaluation metric on RM itself, as we see in section 7.1; and subsequently (2) RM evaluations relying heavily on extrinsic RLHF evaluations. Because RLHF evaluations often rely on human annotations (Wu et al., 2023; Lee et al., 2023), which can be costly and unreliable, they offer limited insights on how we should make research progress on reward modeling. Even though Contrast Instructions do not necessarily assess all capabilities that a RM requires, we hope that it works as a sensible and scalable intrinsic evaluation metric that facilitates future development of better or alternative reward modeling strategies.

Related work

Consistency has been a long-standing topic in NLP research, in previous works, consistency of an NLP mode is defined the invariance of its behavior under meaning-preserving alternations (Ribeiro et al., 2020; Elazar et al., 2021; Goel et al., 2021; Wang et al., 2022b), and several works have explored the consistency in various tasks (Du et al., 2019; Ribeiro et al., 2019; Alberti et al., 2019; Camburu et al., 2020; Asai & Hajishirzi, 2020; Kassner et al., 2021; Chen et al., 2021a; Elazar et al., 2021; Mitchell et al., 2022). In the context of reward modeling, we study the consistency with respect to human preference instead.

Reinforcement Learning from Human Feedback

RLHF (Ouyang et al., 2022; OpenAI, 2023) has merged as a popular technique for aligning LLMs with human preferences (Nakano et al., 2021; Glaese et al., 2022; Bai et al., 2022b; Ouyang et al., 2022; Bai et al., 2022a). The RLHF method involves learning a reward function on human annotations to proxy human preferences, and optimizing language models through reinforcement learning techniques such as Proximal Policy Optimization (PPO) (Schulman et al., 2017). A key implication of RLHF research is to align LLMs with helpful, honest, and harmless human feedback (Askell et al., 2021), (Glaese et al., 2022).

Conclusion

Through the lens of Contrast Instructions, we uncover and study the phenomena of reward inconsistency in reward modeling for RLHF. While our study suggests one perspective and direction on improving RM for RLHF, the question of “what type of reward modeling would most benefit RLHF” remains wide open. We hope that this paper’s findings would facilitate future research and evaluation on this problem.

References

Supplementary Material

Appendix A Details of Contrast Instructions

Using a human preference dataset, we have divided it into training, development, and testing sets. The reward model is trained on the training set and ceases training once it attains optimal performance on the development set. Subsequently, it is evaluated on the test set. Our Contrast Instructions are built upon the test set in each benchmark. To ensure the retrieved instruction differs from the original one, we establish a similarity threshold range (e.g., [0.8,0.9][0.8,0.9]). Only instructions falling within this similarity range are retrieved.

Appendix B Full results on Contrast Instructions

Appendix C Configurations of RL training

There are three models in the RL training stage: the SFT model, the reward model, and the policy model. For two groups of experiments, one uses the original RM (inconsistent), and the other one uses RM equipped with ConvexDA. For the SFT model, both of the groups use the same SFT model, which is fine-tuned on StackExchange. We employ the Low-Rank Adaptor (LoRA) technique (Hu et al., 2021) for training the reward model, with the same configuration in StackLLaMa (Beeching et al., 2023). The training is conducted in an int8 style, due to our computational limits. The learning rate is set to 1.4e−51.4e-5, and we utilize the Adafactor optimizer. The default learning rate scheduler type is set to ‘linear’. The initial KL penalty coefficient is set as 0.2, and an adaptive KL control is used, with a linear scheduler. The pretraining gradient coefficient γ\gamma is set to 0 for our experiments.

Appendix D Consistency of reward modeling when facing adversarial attacks

Adversarial attacks are independent of access to the RM’s training data. Within the context of an adversarial attack process, there are principally two actors: the victim (reward model) Rθ\mathcal{R}_{\theta} and the attack algorithm A\mathcal{A}. (1) Victim: Presented with a human-preference benchmark consisting of benign sentence pairs (those without adversarial perturbations), the victim reward model Rθ\mathcal{R}_{\theta} is trained on these pairs to differentiate the human-preferred response corresponding to the instruction II. (2) Attack algorithm: For a benign sentence pair (rA,rB){(r_{A},r_{B})} with the correct label yAy_{A}, the text attack A\mathcal{A} generates an adversarial sentence rA∗=A(rA){r_{A}}^{*}=\mathcal{A}\left({r_{A}}\right) by adding subtle textual perturbations to rAr_{A}. The goal of these perturbations is twofold: (1) to cause the reward model to issue an incorrect prediction yBy_{B}, and (2) to ensure that the semantics between rAr_{A} and rA∗{r_{A}}^{*} remain closely aligned.

Adversarial consistency refers to a model’s resilience against perturbations generated by adversarial attacks, which try to modify the texts with imperceptible perturbations. Coarsely, these attacks modify the text data at character level (Belinkov & Bisk, 2018; Eger et al., 2019; He et al., 2021), word level (Alzantot et al., 2018; Zhang et al., 2021; Wang et al., 2022a) or sentence level (Jia & Liang, 2017; Ribeiro et al., 2018; Zhang et al., 2019). Basically, the defense methods against adversarial attacks, which can be categorized into three paradigms: (1) model-enhancement-based (Le et al., 2021; Wang et al., 2021; Shen et al., 2022b), (2) certified-robustness-based (Huang et al., 2019; Jia et al., 2019), and (3) detection-based (Mozes et al., 2021; Le et al., 2021; Shen et al., 2023).

We employ a range of existing textual attacks to assess the consistency of the Reward Model (RM). Our chosen attacks encompass three distinct levels of complexity, ranging from straightforward character manipulations to intricate word-level perturbations. For character-level manipulations, we employ methods such as VIPER (Eger et al., 2019) and DeepWordBug (Gao et al., 2018). At an intermediate, word-level complexity, we leverage techniques such as PWWS (Ren et al., 2019), Genetic Attack (GA) (Alzantot et al., 2018), and TextFooler (Jin et al., 2020). Additionally, we introduce a simplistic word-level adversarial method, designated as Vanilla Attack (VA), which exclusively utilizes word-level data augmentation for adversarial perturbations, eschewing additional algorithmic complexity. A detailed description of the Vanilla Attack (VA) is shown in algorithm 1.

D.2 Evaluation

Adversarial attacks originate from an iterative accumulation of adversarial perturbations. Accordingly, we introduce two distinct metrics to encapsulate the model’s response to the incorporation of each successive perturbation: adversarial accuracy, which refers to the RM accuracy on adversarial data, denoted as P\mathcal{P}. These metrics are engineered to illustrate the consistency of RM at both the reward score tier and the performance tier. They are formally defined as follows:

where ϵi\epsilon_{i} is the perturbation generated by attack in the i-th iteration.

D.3 Results

The results are shown in Figure 4. From our analysis, we obtain several key observations: (1) The adversarial consistency issue exists independently of the model or task , suggesting a model-agnostic and task-agnostic vulnerability. Empirically, every model from GPT2-0.1B to LLaMa-7B exhibited vulnerability to adversarial attacks. Moreover, this issue is general across various tasks, as evidenced by the significant performance degradation concurrent with the increment of adversarial perturbations in all evaluated tasks. (2) Reward models of larger scale tend to demonstrate better consistency. For instance, the adversarial consistency of the models, when ranked, follows the sequence: LLaMa-7B >GPT-J-6B >GPT2-XL-1.5B >GPT2-0.1B, which aligns with their size ranking. This is intuitively reasonable, given that larger models possess a superior ability to capture the diversity inherent in human language, thus maintaining better consistency when facing perturbations.

Appendix E Consistency of reward modeling when facing backdoor attacks

Backdoor attack requires access to the training data of RM. Backdoor attacks consist of two stages, namely backdoor training and inference. In backdoor training, the attacker first crafts some poisoned training samples (rA,rB∗,yB)∈D∗\left(r_{A},r_{B}^{*},y_{B}\right)\in\mathcal{D}^{*} by modifying benign training samples (rA,rB,yA)∈D(r_{A},r_{B},y_{A})\in\mathcal{D}, where rB∗r_{B}^{*} is the trigger-embedded input generated from rBr_{B},yBy_{B} is the adversary-specified target label, D∗\mathcal{D}^{*} is the set of poisoned samples, and D\mathcal{D} is the set of benign training samples. Then the poisoned training samples are mixed with the benign ones to form the backdoor training set Db=D∗∪D\mathcal{D}_{b}=\mathcal{D}^{*}\cup\mathcal{D}, which is used to train a backdoored reward model Rθ∗\mathcal{R}_{\theta^{*}}. During backdoor inference, the backdoored model can correctly classify benign test samples: Rθ∗(rA,rB)=yA\mathcal{R}_{\theta^{*}}\left(r_{A},r_{B}\right)=y_{A}, but would classify the trigger-embedded inputs as the target label: Rθ∗(rA,rB∗)=yB\mathcal{R}_{\theta^{*}}\left(r_{A},r_{B}^{*}\right)=y_{B}.

Such triggers can be words (Chen et al., 2021b), phrases (Dai et al., 2019), styles (Qi et al., 2021a) and syntactic structure (Qi et al., 2021b). Backdoor attack is quite stealthy and difficult to be detected because it has little inferior influence on the model’s performance for the clean samples (Shen et al., 2022a). Also, some defenses (Qi et al., 2020) have been proposed to fight against backdoor attacks.

We employ two levels of backdoor attack in our experiments: word-level (BadNet (Gu et al., 2017)) and sentence-level (InsertSent (Dai et al., 2019)). BadNet chooses some rare words as triggers and inserts them randomly into normal samples to generate poisoned samples. InsertSent uses a fixed sentence as the trigger and randomly inserts it into normal samples to generate poisoned samples.

Specifically, for BadNet, we add the word Good! as a word-level trigger appending before each backdoored sentence. For InsertSent, we add the sentence That is a good question! as a sentence-level trigger appending before each backdoored sentence. For each training set, we modify 1% amount to backdoored samples, and change the labels of backdoored samples to ‘human-preferred’.

E.2 Evaluation

We adopt two metrics to evaluate the effectiveness of a backdoor attack: (1) ASR: the classification accuracy on the poisoned test set, which is constructed by poisoning the test samples that are not labeled the target label. This metric reflects the effectiveness of backdoor attacks.; (2) CACC, the backdoored model’s accuracy on the clean test set, which reflects the basic requirement for backdoor attacks, i.e., ensuring the victim model behaves normally on benign inputs.

E.3 Results

The results of our experiments are presented in Table 9. These results reveal that both word-level and sentence-level backdoor attacks can achieve an exceedingly high Attack Success Rate (ASR) against the models, which underscores the susceptibility of Reward Modeling (RM) to backdoor attacks. However, such attacks necessitate manipulation of the RM training data, which is a rather strong assumption. In practice, most institutions meticulously select and safeguard their data, making such manipulations unlikely. Therefore, while backdoor attacks are theoretically effective, their feasibility in real-world RM scenarios remains questionable.

Appendix F Ablation studies of our methods

For every sentence, we juxtapose ConvexDA with standard data augmentation baselines in terms of performance on Contrast Instructions, performance on the original test set, and efficiency. In particular, we choose the WMT19 dataset as our benchmark. For standard data augmentation, we incorporate NN augmented samples for each training sentence, setting NN to 3 and 5, respectively.

F.2 Analyses of RewardFusion

A hyper-parameter in RewardFusion is the threshold of retrieval similarity δ\delta. This part investigates the effect of δ\delta on the performance on original test set and Contrast Instructions, and the results are shown in Figure 9. Among all the choices, δ=0.95\delta=0.95 is the best choice.

Appendix G Case comparisons between RLHF models guided by a consistent and inconsistent RM

Here we demonstrate a few randomly sampled questions and compare the responses from RLHF models trained on the consistent vs. standard RM respectively.