Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
Shicong Cen, Jincheng Mei, Katayoon Goshvadi, Hanjun Dai, Tong Yang, Sherry Yang, Dale Schuurmans, Yuejie Chi, Bo Dai
Introduction
Fine-tuning large language models (LLMs) by reinforcement learning from human feedback (RLHF) (Ziegler et al.,, 2019) has been shown to significantly improve the helpfulness, truthfulness and controllability of LLMs, as illustrated by InstructGPT (Ouyang et al.,, 2022) and many follow-ups. Roughly speaking, there are two critical components of RLHF: (1) reward modeling, which maps human preference rankings into a quantitative reward function that can guide policy improvement; and (2) RL fine-tuning, which seeks to adjust LLM output to align with human preferences by leveraging the learned reward function, i.e., increasing the probability of preferred answers and decreasing the probability of unfavored answers.
Evidently, the curation of preference data is instrumental in the performance of RLHF, which is commonly modeled as pairwise comparisons from a Bradley-Terry ranking model (Bradley and Terry,, 1952). In particular, given a query , human annotators choose a preferred answer from two candidate answers and generated by an LLM. Despite the simple form, collecting large-scale and high-quality preference data can be expensive and time-consuming. Depending on the availability of preference data, two paradigms of RLHF are considered: (1) offline RLHF, where only a pre-collected preference dataset is available, possibly generated from a pre-trained LLM after supervised fine-tuning (SFT); and (2) online RLHF, where additional preference data can be collected adaptively to improve alignment. While initial work on RLHF focused on the offline setting, the online setting has also begun to receive considerable attention, as even a small amount of additional preference data has been shown to greatly boost performance.
There has been significant work on the theoretical underpinnings of RLHF that seeks to uncover algorithmic improvements. Notably, while the original RLHF pipeline decouples reward modeling from RL fine-tuning, direct preference optimization (DPO) (Rafailov et al.,, 2023) integrates these as a single step in the offline setting, leveraging a closed-form solution for the optimal policy in the RL fine-tuning phase. This has led to a welcome simplification of the RLHF pipeline, allowing direct optimization of the policy (i.e., the LLM) from preference data.
Nevertheless, significant challenges remain in RLHF, particularly concerning how to incorporate estimates of reward uncertainty in direct preference optimization when parameterizing policies with large-scale neural networks — such as LLMs — in a theoretically and practically effective manner. In standard reinforcement learning (RL), managing uncertainty when an agent interacts with an environment is a critical aspect in achieving near-optimal performance (Sutton and Barto,, 2018), when using methods that range from policy-based (Schulman et al.,, 2017; Xiao et al.,, 2021), value-based (Mnih et al.,, 2015; Kumar et al.,, 2020), and actor-critic methods (Mnih et al.,, 2016). One dominant approach in the bandit setting, for example, is to construct confidence intervals of the reward estimates, then acting according to the upper and lower confidence bounds — following the principles of optimism and pessimism in the online and offline settings respectively (Lattimore and Szepesvári,, 2020; Lai et al.,, 1985; Rashidinejad et al.,, 2022).
Despite the fact that uncertainty estimation is even more critical in RLHF, due to the coarse nature of preference data, effective implementations of theoretically justified optimistic and pessimistic principles have yet to be developed in the RLHF literature. For example, existing online preference alignment methods, such as Nash-MD (Munos et al.,, 2023) and OAIF (Guo et al.,, 2024), do not incorporate exploration; similarly, pessimism is also not implemented in offline preference alignment methods, such as DPO (Rafailov et al.,, 2023) and IPO (Azar et al.,, 2024). A key reason for these omissions is that it is extremely difficult to construct confidence intervals for arbitrary neural networks (Gawlikowski et al.,, 2021), let alone LLMs. Since optimism for online exploration and pessimism for offline RL both require uncertainty estimation, and given the difficulty of conducting uncertainty estimation for large-scale neural networks, a natural and important question arises:
Can we implement the optimistic/pessimistic principles under uncertainty in a practically efficient manner for online/offline preference alignment in LLMs while retaining theoretical guarantees?
In this paper, we provide affirmative answer to the question. Our major contributions are as follows.
We propose value-incentivized preference optimization (VPO) for both online and offline RLHF, a unified algorithmic framework that directly optimizes the LLM policy with the optimistic/pessimistic principles under uncertainty. Avoiding explicit uncertainty estimation, VPO regularizes maximum likelihood estimation of the reward function toward (resp. against) responses that lead to the highest value in the online (resp. offline) setting, hence implementing optimism (resp. pessimism). Theoretical regret guarantees of VPO are developed for both online and offline RLHF, matching their corresponding rates in the standard RL literature with explicit uncertainty estimation.
In addition, VPO reveals the critical role of reward calibration, where the shift ambiguity of the reward model inherent in the Bradley-Terry model (Bradley and Terry,, 1952) can be exploited to implement additional behavior regularization (Pal et al.,, 2024; Ethayarajh et al.,, 2024). This allows VPO to provide a theoretical foundation for popular conservative offline RL methods (e.g., (Kumar et al.,, 2020)), as well as regularized RLHF methods (e.g., DPOP (Pal et al.,, 2024)).
VPO admits a practically-implementable form suitable for RLHF on LLMs, and more generally, deep-learning architectures. We conduct extensive experimental studies using TL;DR and ARC-Challenge tasks in online and offline settings with optimistic and pessimistic bias, respectively. The results demonstrate improved empirical performance.
2 Related work
Since the introduction of the original RLHF framework, there have been many proposed simplifications of the preference alignment procedure and attempts to improve performance, including SLiC (Zhao et al.,, 2023), GSHF (Xiong et al.,, 2023), DPO (Rafailov et al.,, 2023), and its variants, such as Nash-MD (Munos et al.,, 2023), IPO (Azar et al.,, 2024), OAIF (Guo et al.,, 2024), SPO (Swamy et al.,, 2024), GPO (Tang et al.,, 2024), and DPOP (Pal et al.,, 2024). These methods can roughly be grouped into online and offline variants, depending on whether preference data is collected before training (offline) or by using the current policy during training (online).
In offline preference alignment, identity preference optimization (IPO, (Azar et al.,, 2024)) argues that it is problematic to use the Bradley-Terry model in DPO to convert pairwise preferences into pointwise reward values, and proposes an alternative objective function to bypass the use of the Bradley-Terry model. DPO-Positive (DPOP, (Pal et al.,, 2024)) observes a failure mode of DPO that the standard DPO loss can reduce the model’s likelihood on preferred answers, and proposes to add a regularization term to the DPO objective to avoid such a failure mode. On the other hand, online AI feedback (OAIF, (Guo et al.,, 2024)) proposes an online version of DPO, where online preference data from LLM annotators is used to evaluate and update the current LLM policy in an iterative manner. Iterative reasoning preference optimization (Iterative RPO, (Yuanzhe Pang et al.,, 2024)) proposes to add an additional negative log-likelihood term in the DPO loss to improve performances on reasoning tasks. Finally, (Chang et al.,, 2024) proposes to reuse the offline preference data via reset.
Uncertainty estimation in RL.
The principles of optimism and pessimism are typically implemented via constructing confidence intervals or posterior sampling, which have been demonstrated to be provably efficient in tabular settings (Jin et al.,, 2018; Shi et al.,, 2022). Yet, these approaches have had limited success in conjunction with deep learning architectures (Gawlikowski et al.,, 2021), and many empirical heuristics in turn lack theoretical validation (Kumar et al.,, 2020). VPO draws inspiration from reward-biased exploration (Kumar and Becker,, 1982; Liu et al.,, 2020; Hung et al.,, 2021; Mete et al.,, 2021) in the standard online RL literature, but significantly broadens its scope to the offline setting and RLHF for the first time.
Preliminaries
In RLHF, a language model is described by a policy , which generates an answer given prompt according to the conditional probability distribution . The standard RLHF process consists of four stages: supervised fine-tuning (SFT), preference data generation, reward modeling, and RL fine-tuning. In the SFT stage, a language model is obtained by fine-tuning a pre-trained LLM with supervised learning. The remaining stages continue training by leveraging the preference data, which we elaborate below.
An oracle (e.g., a human labeler or a scoring model) evaluates the quality of two answers and given prompt and reveals its preference. A widely used approach for modelling the probability of pairwise preferences is the Bradley–Terry model (Bradley and Terry,, 1952):
Given a preference dataset composed of independent samples, the reward function can be estimated by maximum likelihood estimation (MLE):
RL fine-tuning.
Given a reward model , we seek to fine-tune the policy to achieve an ideal balance between the expected reward and its distance from an initial policy , which is typically the same as . This is achieved by maximizing the KL-regularized value function , defined as
where \mathsf{KL}\big{(}{{\pi_{1}}\,\|\,{\pi_{2}}}\big{)} is the KL divergence from to , and is a regularization parameter. Consequently, the RL fine-tuned policy with respect to the reward satisfies
which admits a closed-form solution (Rafailov et al.,, 2023), i.e.,
Here, is a normalization factor given by
Direct preference optimization.
The closed-form solution (6) allows us to write the reward function in turn as
Plugging the above equation into the reward MLE (2), we obtain the seminal formulation of direct preference optimization (DPO) over the policy space (Rafailov et al.,, 2023),
which avoids explicitly learning the reward model.
Value-Incentivized Preference Optimization
A major caveat of the standard RLHF framework concerns the lack of accounting for reward uncertainty, which is known to be indispensable in the success of standard RL paradigms in both online and offline settings (Cesa-Bianchi et al.,, 2017; Rashidinejad et al.,, 2022). This motivates us to investigate a principled mechanism that be easily integrated into the RLHF pipeline, while bypassing the difficulties of explicit uncertainty estimation in LLMs.
In view of the sub-optimality of naive MLE for reward estimation (Cesa-Bianchi et al.,, 2017; Rashidinejad et al.,, 2022), and motivated by the effectiveness of reward-biased MLE in online RL (Kumar and Becker,, 1982; Liu et al.,, 2020), we propose to regularize the reward estimate via
We assume that , where
Here, is the prompt distribution and is a fixed calibration distribution independent of the algorithm.
The proposed regularized MLE of the Bradley-Terry model (2) appends a bias term to the negative likelihood
incentivizing the algorithm to favor (resp. avoid) reward models with higher value in the online (resp. offline) setting. Here, is a constant controlling the strength of regularization, and is set to in the online setting and in the offline setting.
At first glance, the objective function for VPO (12) does not immediately imply a computationally-efficient algorithm due to the presence of . However, by exploiting the same closed-form solution for the optimal policy given the reward in (6), and the reward representation inferred from the policy vai (8), we can explicitly express as
where the second step follows because the bracketed term is independent of (c.f. (6)) and the last step follows from (11) whenever . Given this key ingredient, we can then rewrite (12) to directly optimize the LLM policy, in a flavor similar to DPO, as
where we drop the constraint on , since for any policy there exists such that .
In what follows, we elaborate the development of VPO in both the online and offline settings with corresponding theoretical guarantees under linear function approximation.
2 Online RLHF: algorithm and theory
New preference data generation. We sample a new prompt and two answers , query the preference oracle and append to the preference dataset.
Reward learning. We train a reward model with preference data by minimizing the regularized negative log-likelihood, i.e.,
Policy learning. This step trains the policy by solving the RL fine-tuning problem:
We summarize the detailed procedure in Algorithm 1.
Encouragingly, VPO admits appealing theoretical guarantees under function approximation. For simplicity, we restrict attention to linear approximation of the reward model.
Under Assumption 1 and 2, it is sufficient to focus on where
The next theorem demonstrates that Algorithm 1 achieves cumulative regret under mild assumptions. The proof is provided in Appendix A.
Under Assumptions 1 and 2, let denote the corresponding reward model for . Assume that and for some . Then with probability we have
with and .
Theorem 1 shows that VPO achieves the same regret for online RLHF as its counterparts in standard contextual bandits with scalar rewards and using UCB for exploration (Lattimore and Szepesvári,, 2020).
The analysis naturally extends to allowing mini-batch samples of size in every iteration, yielding an improved regret bound scaled by and scaled by .
3 Offline RLHF: algorithm and theory
In offline RLHF, a fixed offline preference dataset is collected , where , are sampled from a behavior policy , such as from SFT. The proposed VPO for offline RLHF consists of one pass through the reward and policy learning phases, i.e.,
which discourages over-optimization of the reward function given the limited offline preference data. In the same vein as deriving (1), and by leveraging (13), we obtain the direct policy update rule:
We summarize the detailed procedure in Algorithm 2. When is set to , the regularization term becomes the KL divergence between and , which is reminiscent of a popular choice in offline RL practice (Kumar et al.,, 2020). Another heuristic choice is to set to the marginalized positive answer distribution from the dataset, i.e., , which leads to a similar objective in (Pal et al.,, 2024).
We first illustrate that VPO indeed executes the principle of pessimism in a complementary manner to the standard approach of pessimism, which finds a policy that maximizes the worst-case value function over a confidence set. In particular, this strategy (Uehara and Sun,, 2021) obtains a policy by solving
As such, the policy obtained by VPO can be equivalently written as
Theoretical analysis.
The next theorem establishes the sub-optimality gap of VPO with linear function approximation under mild assumptions. The proof is given in Appendix B.
Under Assumptions 1 and 2, let denote the corresponding reward model for . Assume that and for some . Let and . With probability , we have
where is the feature sample covariance matrix, , C_{1}=\exp(C)\Big{(}\sqrt{{{d+\log(1/\delta)}}}+\kappa_{\mathcal{D}}\Big{)}+C and . Here,
Experiments
In this section, we evaluate the proposed VPO on both synthetic multi-armed bandit (MAB), and RLHF for LLMs, in online and offline settings.
We evaluate the proposed methods on a synthetic dataset of size and . We set , where with sampled i.i.d. from . The ground truth reward is randomly generated i.i.d. according to . We approximately solve the optimization problems by performing AdamW optimization steps with learning rate and weight decay rate in every iteration for the online setting and steps for the offline setting.
We plot the average results over 10 independent runs in Figure 1. As demonstrated in the left panel of Figure 1, an appropriate choice of allows our method to outperform the model-based MAB with MLE baseline in the long-term performance of cumulative regret, at the cost of slightly increased cumulative regret in the first 100 iterations. This highlights the effectiveness of the VPO in achieving more principled exploration-exploitation trade-off. For the offline setting, the right panel of Figure 1 demonstrates that the performance of both MLE-MAB and VPO improves as the number of offline data increases. However, VPO achieves a consistently lower sub-optimality gap compared with that of MLE-MAB.
2 RLHF for LLMs
We further evaluate the pessimistic/optimistic VPO for LLMs in offline and online setting, respectively. In both settings, the proposed VPO demonstrates strong performances over the baselines.
In this setting, we test pessimistic VPO on ARC-Challenge task (Clark et al.,, 2018), which contains multiple-choices questions from multiple science subjects. We evaluate the performances on the ARC-Challenge test set, which contains questions. The data set only provides ground truth answer for each question. To construct the preference pairs and their labels, for each correct response in the training split, we create three pairs of comparison between the correct answer and each incorrect answer.
We emphasize that our goal is to evaluate the RLHF algorithm designs for LLMs, rather than pushing LLM towards state-of-the-art performance. To demonstrate the advantages of the proposed VPO, we conduct comparison with several offline RLHF baselines (DPO (Rafailov et al.,, 2023) and IPO (Azar et al.,, 2024)) on several LLMs, including Llama2-7b-chat, Llama2-13b-chat (Touvron et al.,, 2023) and Flan-T5-xl (Chung et al.,, 2022). For fair comparison, we keep all the experiment settings and prompts the same for every RLHF algorithm. We did not apply any additional chain-of-thought reasoning to avoid compounding factors affecting the RLHF performances. We tuned the hyperparameters for both the proposed VPO and the baselines on the validation set to achieve their best performances. For detailed hyperparameters setup, please refer to Appendix C.
The performances are illustrated in Figure 2. As we can see, the proposed VPO method demonstrates significantly better performance over the existing baselines on the three models, verifying the benefits across different models. In particular, the performance benefit becomes more evident for larger models. Another important observation is that the proposed VPO method is more robust to over-optimization (Gao et al.,, 2023). In the experiment, the performances of DPO significantly drops after iterations, and the longer DPO is trained, the worse it performs. In contrast, VPO consistently maintains the performances, avoiding the overoptimization issue and justifying the implicit robustness of pessimism as we revealed in (23).
Online setting.
In this setting, we evaluate the performance of VPO on the TL;DR task (Stiennon et al.,, 2020). We prepare the prompts dataset by extracting the input prompts from the preference data. Recall we are evaluating the algorithm performance in online setting, we only compare to the online RLHF baselines (Guo et al.,, 2024) for fairness. We adopt PaLM2 (Anil et al.,, 2023) as the language model and also the LLM annotator. We conduct VPO and online DPO to the same PaLM2-XXS as the policy, which is initialized by supervised finetuning, denoted as SFT model. We exploit another PaLM2-XS model as the LLM annotator to provide online feedbacks. Similar to (Guo et al.,, 2024), we use Detailed 0-shot prompt from Lee et al., (2023). The prompts we used and how we get preference scores are detailed in Appendix C. We emphasize our algorithm is agnostic to human or AI feedback.
As a sanity check, we track the win rate of VPO and online DPO against the SFT baseline on TL;DR during training in Figure 4.2. For ablation purpose, we varies the exploration weight in the optimistic VPO. One significant observation is that although all the online RLHF algorithms follow the increase trend, the win-rate against SFT of the optimistic VPO has larger oscillation, comparing to online DPO. And the oscillation reduces, with diminishing. Our conjecture is that this behavior is encouraged by the optimistic term in VPO, for collecting more unexplored data, which may delay the learning due to the diversity in data. However, as the learning proceeds, the proposed VPO outperforms the competitors, because of the coverage of the collected data.
To demonstrate the advantages of optimistic VPO in online setting more directly, we evaluate the win/tie/loss rate against online DPO head-to-head, as shown in Figure 4.2. This clearly shows that the optimistic VPO achieves better performances with larger exploration preference, and thus, consolidates our conclusion that i), the simple value-incentivized term makes the exploration practical without uncerntainty estimation; and ii), exploration is potentially beneficial for better model.
Conclusion and Discussion
In this work, we develop a unified approach to achieving principled optimism and pessimism in online and offline RLHF, which enables a practical computation scheme by incorporating uncertainty estimation implicitly within reward-biased maximum likelihood estimation. Theoretical analysis indicates that the proposed methods mirror the guarantees of their standard RL counterparts, which is furthermore corroborated by numerical results. Important future directions include investigating adaptive rules for selecting without prior information and more refined analysis on the choice of . This work also hints at a general methodology of designing practical algorithms with principled optimism/pessimism under more general RL setups.
Acknowledgement
The work of S. Cen and Y. Chi is supported in part by the grants ONR N00014-19-1-2404, NSF DMS-2134080, CCF-2106778 and CNS-2148212. S. Cen is also gratefully supported by Wei Shen and Xuehong Zhang Presidential Fellowship and JP Morgan AI Research Fellowship.
References
Appendix A Analysis for the online setting
For ease of presentation, we assume that is finite, i.e., . The general case can be directly obtained using a covering number argument, which we refer to (Liu et al.,, 2024; Jin et al.,, 2022) for interested readers.
We start by decomposing the regret into two parts:
The following lemma is adapted from (Liu et al.,, 2024, Proposition 5.3), whose proof is deferred to Appendix A.2.
Let . With probability , we have
Putting the above inequalities together, it holds with probability that
Step 2: breaking down term (ii) with the elliptical potential lemma.
The linear function approximation form (18) allows us to write
for some . We begin by decomposing term (ii) as
where is an indicator function of event . To proceed, we recall the elliptical potential lemma for controlling the cumulative sum of .
Here, (i) is due to Cauchy–Schwarz inequality, (ii) is due to for , and (iii) results from Young’s inequality. We leave the constant to be determined later.
The second term of (33) can be bounded by
where the first inequality follows from and since .
Putting (33), (35) and (36) together, we arrive at
Step 3: continuing bounding term (ii).
It boils down to control \big{\langle}{W(r^{(t)}),X(r^{(s)})}\big{\rangle}^{2}. We have
where . Therefore,
Recall from (6) that . It follows that (see e.g., (Cen et al.,, 2022, Appendix A.2)), and hence . To proceed, we demonstrate in the following lemma that can be upper bounded by the corresponding Hellinger distance, whose proof is deferred to Appendix A.3.
Assume bounded reward , . We have
where we denote . Plugging the above bound into (37), we get
Step 4: finishing up.
Combining (25), (29) and (40), with probability we have
as long as . Setting , , and , we arrive at
A.2 Proof of Lemma 2
To proceed, we recall a useful martingale exponential inequality.
Let be a sequence of real-valued random variables adapted to filtration . It holds with probability such that for any ,
Applying the above lemma to along with the filtration with given by the -algebra of , we conclude that it holds with probability that
where the last step results from the inequality for all . To proceed, note that
where we denote by the summation over different comparison results. Plugging the above equality into (44) completes the proof.
A.3 Proof of Lemma 4
for some between and . Since , we have
Appendix B Analysis for the offline setting
Therefore, for any , by convexity of the objective function we have
B.2 Proof of Theorem 2
We decompose the sub-optimality gap of by
where the last line is due to according to the definition of (c.f. (20)). We proceed to bound the two terms separately. Here we have written for notational simplicity. In addition, we denote the MLE estimate by .
By the definition of (cf. (4)), it follows that term (i) in (46) can be further decomposed as
where .
To continue, we recall a useful lemma from (Zhu et al.,, 2023).
For any and , with probability at least ,
for all such that
The first term of (47) can be bounded with Lemma 6 as
Step 2: bounding term (ib).
With linear constraint (19), by KKT condition we have
The penultimate step results from , which ensures
Therefore, the second term of (47) can be bounded as
Step 3: bounding term (ii).
where the last step is due to (51). On the other hand, with probability we have (Zhan et al.,, 2023, Lemma 1):
Step 4: putting things together.
Combining (46) (53), (54), with probability we have
Here we have set and . We conclude by bounding as
Appendix C Experimental details
For the offline setting experiments, we adopt instruction tuned models, Llama2-13b-chat and Flan-T5-xl as the base models. To prompt these models, we prepend the question with
What is the choice to the following Question? Only provide the choice by providing a single letter.
The question is structured in a way that the multiple choices are shown as alphabets (letters) within parenthesis.
We set to the empirical distribution of the ground truth answer which is known to us. Based on preliminary experiments, we set as 0.1 in DPO and as 1.0 in IPO. For VPO, we experiment with moving from 0.01 to 10, choosing 1 for the reported results for Flan-T5-xl results. For experiments on Llama2-13b-chat, we also set to 1.
For both models, we train the base models with different algorithms DPO, VPO and IPO for 3000 steps and report the accuracy of the performance on the ARC-challenge test data set after every 500 steps. The training for Llama2-13b-chat model on 128 TPU-v4 takes around 2hrs and for Flan-T5-xl on 64 TPU-v3 takes 1 hour.
C.2 Online Setting
The prompt used for generating AI feedback (and rating) for TL;DR summarization is identical to (Guo et al.,, 2024). We set to the empirical distribution of the negative answer pairs collected by the policy. We set as 0.1 for the DPO term similar to (Guo et al.,, 2024). Additionally for VPO, we decrease the coefficient exponentially following . We try different values of and report the results for 0.1 and 0.01.
The training of the policy, PaLM2-XXS on 64 TPU-v3 for 5000 steps takes around 12 hours for both online DPO and VPO. We report the win rate percentage against the base SFT model for every 1000 steps using PaLM2-XS judge. We also further conduct side by side comparison of Online DPO and VPO at 5000 step.