Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data

Fahim Tajwar, Anikait Singh, Archit Sharma, Rafael Rafailov, Jeff Schneider, Tengyang Xie, Stefano Ermon, Chelsea Finn, Aviral Kumar

Introduction

Pre-training endows a large language model (LLM) with knowledge about the world. Yet, it does not provide a lever to control responses from these models, especially when we want these solutions to optimize some task-dependent success criteria (e.g., align with human preferences, optimize correctness or compactness). To align LLMs with downstream success criteria, they are then fine-tuned with downstream objectives after pre-training. In this paper, we focus on fine-tuning problems that aim to optimize for binary preferences (from humans or other AI models). A plethora of methods have been proposed for this sort of fine-tuning, including supervised learning on filtered responses (Gulcehre et al., 2023), contrastive training (Rafailov et al., 2023), and on-policy reinforcement learning (RL) (Ouyang et al., 2022) on a reward function extracted from human preferences.

In theory, while all of these methods aim to discover identical optimal policies, achieving this in practice would require full data coverage and infinite computation. These requirements are not met in practice, and hence, the choice of the loss function and the optimization procedure affects performance. However, a lack of a clear understanding of different approaches, coupled with different tradeoffs in implementation, has resulted in substantial confusion: practitioners are unsure as to: (1) whether RL (Ouyang et al., 2022) is required at all, or contrastive approaches (Rafailov et al., 2023; Gheshlaghi Azar et al., 2023), supervised fine-tuning are good enough; and (2) whether preference data should be collected with models in the loop (i.e., in an “on-policy” fashion) or not.

Our goal is to provide clarity on these questions by performing a rigorous study to understand the behavior of existing methods when optimizing for preferences. Our study operates under assumptions typical in preference fine-tuning, including the existence of an underlying ground truth reward function that explains the preference data. We study methods that train an LLM policy to optimize a surrogate loss given by the expected reward under a model of the reward function (learned from preference data) penalized by the KL-divergence between the policy and a reference policy.

To answer the above questions, we develop an analysis framework consisting of didactic bandit problems, synthetic LLM problems, and full-scale LLM problems, constructed out of AlpacaFarm (Dubois et al., 2024) and UltraFeedback (Cui et al., 2023). We then study behaviors of different methods given coverage conditions and geometric relationships in the problem. Our main observation is that algorithms that use on-policy RL in a reward model or attempt to push-down likelihood on certain responses, i.e., utilize a negative gradient term as in contrastive objectives tend to outperform other offline supervised objectives with no on-policy sampling or negative gradient. This is surprising because both on-policy and offline methods still utilize the same data for learning. We also find that using on-policy sampling and negative gradients are especially important when high-reward responses appear in less-likely regions of the reference policy distribution, and provide benefits complementary to each other. In particular, we find that supervised objectives such as Pref-FT and Binary Feed-ME (Dubois et al., 2024) are not able to effectively move probability mass from low reward responses to high-reward responses. Sampling on-policy responses for training, contrastive learning, or employing both on-policy sampling and contrastive training can accomplish this.

We theoretically show that approaches that use on-policy RL or certain variants of contrastive training exhibit “mode-seeking” behavior, resulting in faster accumulation of probability mass on a subset of high-reward responses during learning. This behavior is in contrast to “mode-covering” supervised objectives that attempt to increase likelihood on all high-reward responses, and as a result, are unable to efficiently increase probability mass enough on one subset of high-reward responses. We then compare the behavior of a representative mode-seeking objective, the reverse KL-divergence, with the mode-covering forward KL-divergence to formalize this behavior for categorical distributions. Conceptually, this ability to commit to a certain subset of high-reward responses enables algorithms with on-policy sampling (and optionally, a negative gradient) to perform better than likelihood.

Our work presents several actionable takeaways for practitioners. First, we tie the performance of various methods to geometric conditions on the problem, which can inform practitioners which approach to use. Second, we observe a tradeoff between drawing more on-policy samples and performing more gradient steps with a different policy training objective. Understanding this tradeoff is useful for practitioners since on-policy sampling and training present different computational tradeoffs. Finally, since the performance of fine-tuning is tied to the data composition, we study the effect of conditions on the coverage of the preference data, which could inform data collection.

Related Work

A dominant recipe for fine-tuning LLMs is to run supervised next token prediction (“supervised fine-tuning”) on a dataset of high-quality responses to obtain a good policy initialization. This is followed by fine-tuning on a dataset of human preferences (Casper et al., 2023; Ouyang et al., 2022). This fine-tuning can use on-policy RL methods such as REINFORCE (Sutton et al., 1999) or PPO (Schulman et al., 2017) to maximize the predictions of a reward model obtained from the preference data, regularized with a KL constraint. Another approach (Dubois et al., 2024) performs supervised fine-tuning on the filtered set of preferred completions in the preference dataset. A different family of methods runs supervised learning on preferred responses iteratively such as ReST (Gulcehre et al., 2023), RWR (Hu et al., 2023), and SuperHF (Mukobi et al., 2023). Alternatively, methods such as DPO (Rafailov et al., 2023), IPO (Gheshlaghi Azar et al., 2023), SLiC-HF (Zhao et al., 2023), and KTO (ContextualAI, 2024) learn directly from human preferences, with no explicit reward model. Concurrent work also runs DPO iteratively (Yuan et al., 2024; Chen et al., 2024). These methods come with different tradeoffs necessitating a study to understand their behaviors.

Prior analysis work. To understand the effect of preference fine-tuning, prior work attempts to uncover its effect on network parameters for a certain set of tasks (Jain et al., 2023; Lee et al., 2024). Our analysis is complementary in that it studies conditions when different algorithms perform well, and is applicable to any downstream task. Kirk et al. (2023) study the contribution of RL fine-tuning on generalization to out-of-distribution prompts but this is complementary to our approach. Gao et al. (2022); Coste et al. (2023); Eisenstein et al. (2023) study reward over-optimization to better build reward models, which is complementary to the behavior of the policy optimization approach. Agarwal et al. (2023) develop a recipe that uses the mode-seeking KL divergence for knowledge distillation: this prior work is largely centered in the problem setting of distillation and does not study the optimization behavior of RL, contrastive, or supervised objectives. Perhaps closely related to our work is Singhal et al. (2023), which investigates the interplay between PPO and the composition of preference data, but this analysis is largely concentrated on studying the length bias of RL fine-tuning rather than developing insights into the behavior of fine-tuning algorithms. We do design didactic examples that use rewards dependent on length, but this is solely for analysis.

Concurrently, Ahmadian et al. (2024) show that REINFORCE may simply be enough for preference fine-tuning of LLMs and complex policy optimization methods such as PPO may not be needed. Our conclusions are mostly complementary, though we do observe that PPO is more robust to sample reuse than REINFORCE. Concurrently, Sharma et al. (2024) compares contrastive and supervised fine-tuning on LLM-generated data, but this work does not study the role of coverage or geometric conditions. Nevertheless their conclusions that various approaches perform similarly when the peak in the reward function (i.e., oracle AI preferences) aligns with the likely regions in the data (i.e., responses generated from the same AI model), thus providing evidence to support our findings.

Characterizing And Unifying Preference Fine-Tuning Methods

Typical preference fine-tuning methods use a variety of objectives including RL, maximum likelihood, and contrastive learning. While the huge number of fine-tuning methods inhibits us from empirically analyzing each of them, in this section we characterize several existing methods into different families and subsequently study a representative member from each family.

Given this reward function r∗r^{*}, preference fine-tuning aims to find the optimum of the reward r∗r^{*}. While the ultimate goal of preference fine-tuning is to find the unconstrained optimum of the reward function, in practice, we often replace the reward function with a reward model. Since the reward model is erroneous, we apply KL-constraint to prevent exploitation in the reward model. To align our results with typical preference fine-tuning procedures, we will consider such a KL-constrained reward optimization as our fine-tuning goal:

Reward model training. In order to fine-tune an LLM policy πθ(y∣x)\pi_{\theta}(\mathbf{y}|\mathbf{x}), Equation 3.1 provides a convenient way to learn a reward model either explicitly (i.e., by fitting a parametric reward model rϕ(x,y)r_{\phi}(\mathbf{x},\mathbf{y})) or implicitly (i.e., via direct preference optimization (DPO) Rafailov et al. (2023) or IPO (Gheshlaghi Azar et al., 2023), that re-purposes the log-likelihood log⁡πθ(y∣x)\log\pi_{\theta}(\mathbf{y}|\mathbf{x}) of the policy to represent the reward rθ(x,y)r_{\theta}(\mathbf{x},\mathbf{y})). Explicit reward models are trained using the following classification objective:

where σ\sigma is the logistic function. Contrastive learning objectives (Rafailov et al., 2023; Gheshlaghi Azar et al., 2023) on the other hand repurposes log⁡πθ(y∣x)\log\pi_{\theta}(\mathbf{y}|\mathbf{x}) as the implicit reward rθ(x,y)r_{\theta}(\mathbf{x},\mathbf{y}):

2 Characterizing Fine-Tuning Methods

With a reward model rϕ(x,y)r_{\phi}(\mathbf{x},\mathbf{y}), most fine-tuning approaches attempt to discover the policy πθ(y∣x)\pi_{\theta}(\mathbf{y}|\mathbf{x}) which optimizes Equation 3.2 by using rϕr_{\phi} as a surrogate for r∗r^{*}. Since we cannot empirically investigate all of these methods, we group them into different categories (summary shown in Table 1). In particular, we are interested in whether these methods employ:

on-policy sampling: an explicit sampling of new responses from the policy (e.g., PPO, REINFORCE) or purely learning from offline data (e.g., RWR, DPO, IPO)

on-policy sample reuse: for only those approaches that perform on-policy sampling, whether the approach makes more than one gradient update on a given prompt-response (x,y)(\mathbf{x},\mathbf{y}) pair (e.g., exactly 1 update for REINFORCE, ≥1\geq 1 for PPO, online RWR)

negative gradient: whether the approach explicitly minimizes a loss that attempts to “push-down” likelihood on certain responses by multiplying the gradient of their likelihood with a negative coefficient (e.g., contrastive methods such as DPO; RL methods REINFORCE, PPO)

On-policy RL approaches such as PPO (Schulman et al., 2017) and REINFORCE (Williams, 1992) explicitly sample new responses from the current snapshot of the learned policy, yi∼πθ(⋅∣xi)\mathbf{y}_{i}\sim\pi_{\theta}(\cdot|\mathbf{x}_{i}), score them under the reward model, and perform a policy gradient update on parameters θ\theta, for example:

is the gradient update employed by REINFORCE, where rˉϕ(x,y)\bar{r}_{\phi}(\mathbf{x},\mathbf{y}) corresponds to a normalized estimate of the reward model’s predictions over a batch of samples drawn from the policy. As we discuss in more detail in Section D.1), using a normalized reward estimate instead of directly the raw reward value helps reduce the variance of the policy gradient estimate. High variance gradients slow down convergence and even sometimes lead to sub-optimal solutions in deep RL (Mei et al., 2022).

Due to the use of normalized reward estimates, policy gradient approaches behave distinctly from maximum likelihood supervised learning: a policy gradient update also updates the parameters θ\theta in a direction that attempts to push down likelihood log⁡πθ(y′∣x)\log\pi_{\theta}(\mathbf{y}^{\prime}|\mathbf{x}) for samples y′\mathbf{y}^{\prime} on which normalized reward rˉϕ(x,y′)<0\bar{r}_{\phi}(\mathbf{x},\mathbf{y}^{\prime})<0. This means that on-policy RL also has a form of the “negative gradient”.

PPO differs from REINFORCE because it employs sample reuse in addition to on-policy sampling: unlike REINFORCE which only performs a single gradient update on a response sampled from the current policy, PPO can utilize a response for several policy updates. To prevent making updates on overly off-policy responses, there is a mechanism in place to filter responses by the magnitude of the importance ratio between the current policy πθ(y∣x)\pi_{\theta}(\mathbf{y}|\mathbf{x}) and the data collection policy.

Finally, we also remark that while on-policy methods do generate new rollouts from the policy, these responses are still scored by a reward model (and not the ground truth reward function, i.e., humans). Since reward labels come from a reward model, on-policy preference fine-tuning approaches are instances of offline model-based RL (Yu et al., 2021, 2020; Kidambi et al., 2020) methods that run on-policy rollouts against a learned dynamics and reward model (due to the single step nature of preference fine-tuning, there is no dynamics model).

On-policy supervised approaches such as RAFT (Dong et al., 2023), ReST (Gulcehre et al., 2023), and SuperHF (Mukobi et al., 2023) iteratively minimize a weighted maximum likelihood loss inspired by Peters and Schaal (2007); Korbak et al. (2022). For a given prompt xi\mathbf{x}_{i}, these methods sample NN responses from the model: yi1,⋯ ,yiN∼πθ(⋯∣xi)\mathbf{y}^{1}_{i},\cdots,\mathbf{y}^{N}_{i}\sim\pi_{\theta}(\cdots|\mathbf{x}_{i}), then weight these responses by the exponentiated reward, exp⁡(rϕ(xi,yij)/β)\exp(r_{\phi}(\mathbf{x}_{i},\mathbf{y}_{i}^{j})/\beta) as in the case of reward-weighted regression (RWR) or obtain the subset of KK highest rewarding responses as in the case of ReST or Best-of-N. Finally, these methods train via supervised next-token prediction on these filtered or weighted responses. Given a weighting function, F(xi,yij∣yi0⋯N)F(\mathbf{x}_{i},\mathbf{y}_{i}^{j}|\mathbf{y}_{i}^{0\cdots N}) that maps a response yij\mathbf{y}_{i}^{j} for a given prompt xi\mathbf{x}_{i} to a scalar value conditioned on other responses yik\mathbf{y}_{i}^{k} sampled from the model for the same prompt x\mathbf{x}, these methods maximize:

These algorithms employ sample reuse because they operate in a “batched” online fashion: instead of performing exactly one gradient step on a given model sample; RWR, ReST, and SuperHF run more gradient updates, after which new samples are drawn. However, since these methods only maximize likelihood (i.e., only positive multipliers), there is no negative gradient effect.

Research Questions and Analysis Setup

Our goal is to understand the behaviors of various procedures for fine-tuning language models. As discussed above, typically these methods differ along the use of on-policy sampling (with additional differences pertaining to sample reuse) and the presence of a negative gradient. We build a setup to understand these differences empirically by answering the following questions:

Question 1: When does on-policy sampling improve over offline fine-tuning, even though on-policy samples are annotated by a reward model, which itself is learned from offline data? Is sample reuse useful or harmful for on-policy methods?

Question 2: When does an explicit negative gradient help the discovery of effective policies compared to maximum likelihood approaches such as distilling the Best-of-N policy?

Question 3: Does on-policy sampling offer complementary benefits to negative gradient, resulting in better performance with effective contrastive approaches (e.g., DPO)?

To gain practically useful and actionable insights, we must answer these questions in the context of coverage and geometric relations between the training data, reference policy, and the reward function. These relations affect the shape of the optimally fine-tuned policy and dictate the dynamics of various objectives under consideration. We consider specific conditions and relations that we discuss next.

Understanding the behavior of various approaches as a function of these factors will allow us to better understand the performance of various approaches on downstream fine-tuning in terms of problem geometry [C1] and statistical learning considerations [C2].

2 Tasks and Datasets

We construct a variety of didactic and LLM tasks that allow us to gain intuition for performance of different methods under various scenarios grouped along relationships [C1] and [C2].

Didactic NN-d bandit problems. Equation 3.2 poses preference fine-tuning as a KL-regularized contextual bandit problem over contexts x\mathbf{x}. Therefore, we develop a didactic NN-dimensional contextual bandit problem. We use a set of tokens of size VV of size 100100. The context, x\mathbf{x}, is a single discrete token from VV. A response a\mathbf{a} is a sequence of N=10N=10 discrete tokens from VV. We primarily study the effect of geometric relationship [C1] and assume that the reward function is known exactly, therefore not accounting for the data coverage and training of the reward model. We consider two reward functions that differ in their relative geometry relative to the reference policy, as shown in Figure 2. Specifically, the difference lies in how perfectly the optimum of the reward function aligns with the high-density regions of the reference distribution. The optimum of the reward function R1\mathbf{R}_{1} is located in low likelihood regions of the reference policy, whereas the optimum of R2\mathbf{R}_{2} is roughly aligned with the mode of the reference policy. We hypothesize that on-policy sampling will be crucial to optimize reward function R1\mathbf{R}_{1}, whereas offline or maximum likelihood methods could be sufficient for the optimization of R2\mathbf{R}_{2}.

Synthetic LLM fine-tuning problems. Next, we will generalize our intuitions from bandit problems to the LLM setting. Instead of directly experimenting with human preferences, we first study two synthetic problems that utilize hand-crafted reward functions, which can be approximated via reward models. Access to functional forms of these hand-crafted reward functions will enable us to track the ground-truth objective throughout training to see if our insights about various approaches under condition [C1] will hold even when learning against a reward model. Subsequently, we run this experiment with an altered skewed preference data distribution (see Figure 3) to understand the effect of coverage conditions [C2]. We consider two reward functions: (1) one that minimizes the response length (“Min Length”), analogous to R1\mathbf{R}_{1} in the bandit problem, and (2) that attempts to anchor the response length to a pre-specified target value (“Avg Length”), which lies in the mode of the target distribution. This second condition exhibits similar characteristics to R2\mathbf{R}_{2}. The Skew Length scenario skews the preference data in the Min Length problem scenario.

Full-scale LLM fine-tuning. Finally, we scale up our study to full-scale LLMs, with real preference data. Recent work (Singhal et al., 2023) shows that preference labels are usually biased towards much longer responses, indicating that preference fine-tuning usually admits a geometric relationship where the mode of the reward function is distinct from the mode of human data (and hence, any reference policy). For the majority of our experiments, we use preference datasets from the AlpacaFarm benchmark (Dubois et al., 2024). We also scale up our experiments to UltraChat (Ding et al., 2023), a ∼10\sim 10 times larger dataset with responses from many strong LLMs such as GPT 4 and GPT-3.5.

3 A Generic Fine-Tuning Algorithm Encapsulating All Axes

To systematically analyze the behavior of fine-tuning methods that differ along the axes discussed in Section 3.2, in this section, we introduce a generic algorithm with different hyperparameters associated with each axes. With a generic algorithm of this sort, we will be able to answer our research questions by varying each hyperparameter. Our unified practical algorithm is shown Algorithm 1. While on-policy algorithms perform steps 1 and 2 of on-policy data collection with a reward model, purely offline methods (e.g., DPO and RWR) utilize preference data directly.

To study the impact of on-policy sampling, we vary the extent to which updates are made on data from the current policy. We can control this by two means in Algorithm 1: (1) by varying the total number of samples ∣D∣=BC×C=B|\mathcal{D}|=\frac{B}{C}\times C=B used for a given training iteration assuming the algorithm performs exactly one pass over all this sampled data while keeping the minibatch size MM fixed, and (2) by varying the number TT of gradient steps performed on a given set D\mathcal{D} of on-policy samples (i.e., a larger TT leads to more off-policy updates). In other words, approach (1) will perform more updates using stale data for large values of ∣D∣|\mathcal{D}|; and for small values of ∣D∣|\mathcal{D}|, approach (2) will make more off-policy updates if TT is larger. While both approaches enable us to control how on-policy an algorithm is, approach (1) does not reuse samples (since D\mathcal{D} is large), but approach (2) reuses samples for different number of gradient updates, controlled directly by TT. By studying both approaches for inducing off-policyness, we can isolate the effect of sample reuse on on-policy methods. We also study offline methods with no on-policy sampling, such as DPO, and filtered supervised learning on the preferred response yw\mathbf{y}_{w} in the dataset to understand the role of the negative gradient.

Empirical Analysis Results

In this section, we will present the results of our empirical study to answer our research questions. To answer each question, we will begin by studying the didactic bandit problem with the ground-truth reward function, followed by synthetic and then full-scale LLM fine-tuning problems.

To understand the role of on-policy sampling, we will investigate if on-policy sampling can improve performance for several approaches followed by making conclusions regarding sample reuse.

We first study on-policy sampling as a function of the geometric relationship [C1] in our bandit setting (see Figure 2), with no sampling error. Then, we will extend our conclusions to the LLM setting.

That said, we also note in Figure 4 that the performance degradation with more off-policy updates is substantially milder for R2\mathbf{R}_{2}, indicating that when the peak in the reward function lies in the high likely regions of the reference policy, a higher degree of off-policy updates is tolerable.

Synthetic LLM problems. In this problem setting, we optimize the policy against a reward model, which is learned from preference data. Per Section 4.2, we construct three scenarios that differ along geometric ([C1]) and coverage ([C2]) conditions as depicted in Table 2. The peak of the reward in the Min Length scenario appears in the less likely regions of πref\pi_{\text{ref}}, whereas the peak of the reward function in the Mode Length scenario appears in highly likely regions under πref\pi_{\text{ref}}.

We present our results for one algorithm in detail (in this case, PPO) (Figures 5, 6 and 7) and then present a summary bar chart showing that our conclusions also transfer to other algorithms (such as REINFORCE and RWR) (Figure 8). Extending insights from the bandit problem, in the Min Length scenario, we find that being more on-policy (i.e., a smaller BB) leads to a lower completion length and hence a higher gold reward, despite potential inaccuracies in the proxy reward model that PPO is actually optimizing (Figure 5). Akin to our bandit experiments, we also observe that smaller batch sizes (B=64B=64 and B=128B=128) optimize the proxy reward at a faster rate compared to B=192B=192 and B=256B=256. This indicates that with a significant overlap between the preference data and the reference policy, on-policy sampling still leads to better performance with fewer updates. We also find similar trends across on-policy variants of RWR and REINFORCE, where modulo training instabilities, being more on-policy results in better performance (Figure 8; Min Length).

In the Mode Length scenario, where the preferred response for each preference pair are those that are closest to the average length in the dataset (203), varying the degree of on-policy sampling by adjusting the sampling frequency largely does not affect either the proxy or gold reward for PPO (Figure 6). We make similar observations for other algorithms: Figure 8; Mode Length: different degrees of on-policyness perform similarly, except the more on-policy runs sometimes exhibit instability. This is in agreement with the results from the bandit setting above: when the peak in the reward function lies in highly likely regions under the reference policy, on-policy sampling has minor effect and more off-policy configurations of the algorithm can perform similarly too.

Finally, to evaluate the robustness of these findings under more challenging coverage conditions, we deliberately skew the length distribution in the preference dataset to make it distinct from the reference policy (called Skew Length). Concretely, with a 95% probability, we truncate the length of the response by sampling a length from an exponential distribution, which naturally leads to a shorter completion length. The remaining 5% of samples are drawn from the standard SFT policy to simulate the broader coverage for the preference data. Overall, the resulting data admits a significantly skewed distribution over response lengths, as visualized in Figure 3. Not only does the peak in the reward function now appear in less likely regions of the reference policy, but to succeed, an optimization algorithm must now do the required heavy lifting to shift the probability mass to the low-density regions of the response space that maximize reward.

Our detailed results of running PPO in this setting are shown in Figure 7. In this setting, we still find that more on-policy updates lead to a higher gold reward with PPO. In addition, we also observe much larger gaps in proxy reward values attained at any given gradient step compared to the Min Length scenario, in favor of on-policy sampling. For other algorithms, we also observe strong and clear trends supporting that on-policy sampling with a smaller but frequently sampled batch results in better performance as shown in the summary plot (see Figure 8; Skew Length).

Full-scale LLM problems. Finally, we evaluate if our insights transfer to the full-scale AlpacaFarm setup. We use a Pythia-1.4B model as our reference policy and generate two responses per prompt. We label the preferred and dispreferred responses with a gold reward model of human preferences from AlpacaFarm to construct a preference dataset. Figure 9 shows that our intuitions from the simple bandit and synthetic LLM experiments transfer to this real preference learning task, as making updates on only on-policy samples leads to higher gold reward for both on-policy RWR and REINFORCE.

In the previous section, exactly one gradient step was taken on a given sample and we found that making updates on stale data was not helpful due to off-policy updates. Is there any scenario under which we can still attain good policy performance despite employing off-policy updates? In this section, we will answer this question, and show that it might be possible to learn with off-policy updates for some algorithms if we are allowed to make more than one update on a given sample. Of course, a substantial amount of sample reuse is detrimental since it would lead to more off-policy updates, thus leading to statistical or even propensity overfiting (Swaminathan and Joachims, 2015) for some methods, but it is reasonable to surmise that some amount of sample reuse can help. To study sample reuse, we compare methods when T>1T>1 gradient steps can be made on a given sample.

We study sample reuse for on-policy RWR in the bandit setting in Figure 10. While increasing TT can slow down convergence in general, we note that using a larger value of TT may be better (e.g., T=5T=5 learns faster than T=2T=2; T=10T=10 learns faster than T=7T=7).

Synthetic LLM problems. We also evaluate the effect of sample reuse on synthetic LLM problems. In this case, we study two algorithms PPO and on-policy best-of-N to be able to understand the effect of sample reuse on multiple algorithms. In contrast to the performance degradation with off-policy updates induced due to stale samples in PPO, we find that off-policy updates induced due to sample reuse do not hurt performance (Figure 11; PPO), with even T=8T=8 performing similarly to T=1T=1. On the other hand increasing TT from 11 to 22, i.e., performing two gradient updates on each sample improves the golden reward for best-of-N (Figure 11; Best-of-N) within a given data sampling budget.

Why do PPO and best-of-N respond differently to sample reuse? We believe that this is because PPO employs an off-policy correction, and hence, significantly off-policy samples do not contribute to the gradient, addressing the well-known challenge of propensity overfitting (Swaminathan and Joachims, 2015). This is not the case with on-policy best-of-N, where excessive sample reuse can hurt exploration, because training on old samples with a log-likelihood loss push the current policy to be close to the stale data-generating policy. That said, more than one gradient step can still be useful when presented with a fixed data budget, unless it bottlenecks exploration of high reward regions.

2 Question 2: The Role of Negative Gradient

To understand the role of negative gradient, we will compare contrastive algorithms such as DPO and IPO with maximum likelihood methods such as RWR (or Pref-FT, which attempts to increase the likelihood of the preferred response only) and best-of-N in a fully offline setting, where no new on-policy samples are used. We will also aim to understand the mechanisms behind these methods.

Having seen that using a negative gradient leads to much better performance, we next attempt to understand the mechanism behind this better performance. To do so, we visualize the evolution of the log-likelihoods of the preferred response and the dispreferred response in a held-out dataset as multiple gradient steps are taken on an offline preference optimization loss.

Contrastive training increases the gap between the likelihoods of preferred and dispreferred responses. Perhaps as expected, we find that DPO-style contrastive training is more effective at increasing the gap between the likelihoods of preferred and dispreferred responses compared to offline Pref-FT in several LLM settings: the synthetic LLM settings with Min Length and Skew Length, and full-scale AlpacaFarm and UltraFeedback settings (Figure 15). More concretely, note that the margin for Pref-FT largely converges to 0, whereas offline DPO can enable a larger margin.

We also observe a similar trend in full-scale LLM experiments in Figure 17: we observe a decrease in the log-likelihoods of both the preferred and dispreferred responses throughout training on AlpacaFarm with small 1.4B Pythia policies. However, using a Mistral7B model to train a policy on the UltraFeedback dataset results in an increasing value of log-likelihood of πθ(yw∣x)\pi_{\theta}(\mathbf{y}_{w}|\mathbf{x}) and a decreasing value of πθ(yl∣x)\pi_{\theta}(\mathbf{y}_{l}|\mathbf{x}) when starting from an SFT model on the Ultrachat-200K dataset (same setup as Zephyr (Tunstall et al., 2023)). We believe that these opposite trends are a consequence of the responses in that the UltraFeedback dataset are more semantically distinct from each other, as different responses come from models with different capabilities (e.g., a GPT-4 response is paired with a GPT-3.5 response) such that given enough model capacity, contrastive training can push up likelihoods of πθ(yw∣x)\pi_{\theta}(\mathbf{y}_{w}|\mathbf{x}) while pushing down πθ(yl∣x)\pi_{\theta}(\mathbf{y}_{l}|\mathbf{x}). In contrast, perhaps as expected, running Pref-FT increases the likelihoods of both yw\mathbf{y}_{w} and yl\mathbf{y}_{l} (Figure 17).

3 Question 3: On-Policy Sampling and Negative Gradients are Complementary

Based on our findings that both on-policy sampling and negative gradients are independently effective, we now study if combining them would provide any additional benefits. To understand this, we empirically study a straightforward on-policy variant of DPO/IPO: instead of utilizing the PPO or Best-of-N objective on on-policy samples, for each prompt x\mathbf{x}, we sample NN responses from the policy y1,…,yn∼πθ(.∣x)\mathbf{y}_{1},\ldots,\mathbf{y}_{n}\sim\pi_{\theta}(.|\mathbf{x}), rank them according to a reward model rϕr_{\phi}, and construct preference pairs by taking the higher reward completion as the preferred one and lower reward completion as the dispreferred one. This recipe is similar to concurrent works such as Rosset et al. (2024). Then we calculate the DPO/IPO loss on this preference dataset and update our model accordingly.

Performance on bandit and synthetic LLM problems. Figure 18 shows that the on-policy version of IPO achieves both faster convergence and better performance compared to the offline version, for both R1\mathbf{R}_{1} and R2\mathbf{R}_{2} in the didactic bandit problem. We also ran on-policy DPO in synthetic LLM problems we studied and found it to converge significantly faster and to a better solution than offline DPO, on-policy RL, and on-policy variants of supervised learning approaches as shown in Figure 19. We also find that on-policy versions of contrastive approaches exhibit favorable computational vs wall-clock time tradeoffs compared to purely on-policy RL methods and even offline contrastive methods that may not find as good solutions as their on-policy counterparts (see Appendix B).

Why can on-policy versions of contrastive methods perform better than on-policy RL? We saw in Section 5.2.1 that offline contrastive training with a negative gradient was effective at quickly reorganizing probability mass to high-reward responses covered by the preference data. When combined with on-policy sampling, this behavior results in faster convergence: for any given batch of on-policy data, contrastive training with a negative gradient can quickly reconfigure the policy distribution within the support of the on-policy data obtained thus far (i.e., it provides a stronger, low-variance learning signal). Similarly to how best-of-N + negative gradient outperforms vanilla best-of-N but underperforms DPO in Figure 12, PPO also improves over RWR without a negative gradient term (in the bandit setting this corresponds to a better reward-KL tradeoff in Figure 18 and in the synthetic LLM setting this appears in final performance), but it is still unable to match on-policy DPO in Figure 19. Note that this does not mean that on-policy DPO would always outperform PPO, but that it might be a good choice for users to experiment with on-policy versions of contrastive methods.

Conceptual Unification and Theoretical Analysis

With empirical results showing the benefits of on-policy sampling and negative gradient for preference fine-tuning of LLMs, in this section, we attempt to conceptually understand the benefits by building a mental model. In this section, we will first unify these seemingly distinct notions of on-policy sampling and negative gradient into a unified notion of mode-seeking objectives, and contrast them against mode-covering maximum likelihood objectives. Then, we will contrast the learning dynamics of the reverse KL-divergence, a representative mode-seeking objective against the mode-seeking forward KL-divergence (i.e., the supervised learning loss) to intuitively explain some of our findings.

In this section, we will show that the notion of mode-seeking divergences unifies on-policy sampling and negative gradients for the various objectives we investigated in the paper. Specifically, we show below that several on-policy RL methods that we studied optimize the reverse KL-divergence, and are hence mode-seeking, offline contrastive methods that employ a negative gradient are also mode-seeking, and finally, supervised weighted maximum likelihood approaches (e.g., offline Best-of-N, Pref-FT, Binary FeedMe) are mode-covering. First, we show that on-policy sampling leads to mode-seeking behavior. To do this, we prove that RL and supervised objectives combined on-policy sampling optimize the reverse KL divergence, which is known to be mode-seeking.

On-policy RL and on-policy weighted-likelihood methods optimize a regularized version of a reverse KL-divergence with respect to the optimal policy and are hence mode seeking.

A proof for Lemma 6.1 is shown in Section C.1.1. Next, we show that offline contrastive methods that employ a negative gradient are also mode-seeking. While these approaches do not optimize the reverse KL-divergence, we can still show that the probability mass obtained by minimizing density on negative responses yl\mathbf{y}_{l} gets disproportionately utilized, far more for increasing the probability mass on the “mode” (i.e., highest probability categories under the current policy πθ\pi_{\theta}) compared to other categories. When the offline dataset consists of multiple high-reward categories, this preference to put more probability mass on the mode of the current policy results in mode-seeking behavior, compared to increasing probability mass on all high-reward categories.

Let θt\theta_{t} denote the parameters of the model at a given iteration tt. Consider contrastive approaches that induce a negative gradient under a functional form shown below:

where c1c_{1} and c2c_{2} are non-negative functions that depend on the reward value and the associated samples, yw\mathbf{y}_{w} and yl\mathbf{y}_{l}. In contrast, weighted maximum likelihood without the negative gradient sets c2=0c_{2}=0. Define ωt:=log⁡πθ(yw∣x)−log⁡πθ(yl∣x)\omega_{t}:=\log\pi_{\theta}(\mathbf{y}_{w}|\mathbf{x})-\log\pi_{\theta}(\mathbf{y}_{l}|\mathbf{x}). Then, for all models θ\theta and for all tt, there always exists an appropriate dataset of positive and negative samples D\mathcal{D}, such that:

In addition, if the model class πθ\pi_{\theta} and yl\mathbf{y}_{l} can jointly realize the following gradient alignment condition (note that for any θt\theta_{t}, there always exists a yl\mathbf{y}_{l} that satisfies this condition):

then, we find that the likelihood of positives is larger (and similarly likelihood of negatives is smaller) when c2>0c_{2}>0, i.e., when a negative gradient term is used:

A proof for Lemma 6.2 is provided in Section C.1.2. This result indicates that for appropriate negative responses, a contrastive update accelerates the rate of increase of probability mass on yw\mathbf{y}_{w}, for any model class πθ\pi_{\theta} and reference initialization θ0\theta_{0}, compared to setting c2=0c_{2}=0, which offline weighted maximum likelihood. This corresponds to mode-seeking behavior. The update induced by DPO admits a similar form (see the discussion after Equation 7 in Rafailov et al. (2023)). This theoretical result also corroborates our findings in the experiments in Section 5.2.1 regarding the negative gradient term. The gradient of IPO also admits a similar form (Section C.1.2).

Next, we note that purely offline versions of supervised methods such as RWR, ReST, and BoN, that only maximize weighted likelihood are mode-covering because these objectives can be shown to maximize the forward KL-divergence against the optimal policy (proof in Section C.1.3).

Consider offline supervised methods that maximize weighted log-likelihood:

where F(x,y)≥0F(\mathbf{x},\mathbf{y})\geq 0 is the weight for (x,y)(\mathbf{x},\mathbf{y}). Furthermore, ∑yF(x,y)>0\sum_{\mathbf{y}}F(\mathbf{x},\mathbf{y})>0 (i.e., for every x\mathbf{x}, there exists a response y\mathbf{y} with non-zero F(x,y)F(\mathbf{x},\mathbf{y})). Then these methods optimize a forward KL-divergence.

2 Case Study: Mode-Seeking Reverse KL vs. Mode-Covering Forward KL

Having seen that mode-seeking and mode-covering divergences can unify on-policy sampling and negative gradients, in this section, we perform a theoretical analysis to quantify the behavior of the two representative mode-seeking and mode-covering objectives: reverse KL (mode-seeking) and forward KL (mode-covering) objectives on categorical distributions, parameterized via independent logits. Our goal is to formalize the intuition that a mode-seeking objective can sharpen the probability mass on only certain high-reward regions, thereby leading to aggressive reorganization of probability mass. This helps corroborate our experiments that on-policy sampling in a reward model and offline negative sampling is still useful to quickly align the policy with the target distribution.

Notation and setup. For this result, we will study training a categorical distribution p(x)p(\mathbf{x}) to match the theoretically optimal fine-tuned policy, q(x)q(\mathbf{x}). We assume that p(x)∝exp⁡(f(x))p(\mathbf{x})\propto\exp(f(\mathbf{x})), where each logit f(x)f(\mathbf{x}) is an independent parameter. We train p(x)p(\mathbf{x}) by performing gradient descent, starting from an initial reference distribution p0p_{0} on a fine-tuning loss with gradient descent and a learning rate η\eta. We denote the distribution at step tt of this gradient descent as ptp_{t}. For this analysis it would be helpful to explicitly write out the parameter updates at any iteration tt, induced by forward and reverse KL.

For any given distribution ptp_{t}, with pt(x)=exp⁡(ft(x))p_{t}(\mathbf{x})=\exp(f_{t}(\mathbf{x})), the updates induced by the forward and reverse KL-divergences within one step of gradient descent with a learning rate η\eta are given by:

For a proof of Lemma 6.4, see Section C.2. In principle, upon convergence, both the reverse and forward KL-divergences should find the optimally fine-tuned distribution, q(x)q(\mathbf{x}) in this simple setting. But to understand their behavior in relevant practical situations, we are particularly interested in understanding their behavior at intermediate points during training, when either divergence is not minimized to exactly 0. Insights about intermediate points in training can make useful predictions about practical problems when early stopping is used to prevent overfitting and the loss is rarely 0. Thus, our result below attempts to characterize these objectives at any given iteration tt:

Let pt+1f(x)p^{f}_{t+1}(\mathbf{x}) be the distribution obtained after one gradient step, starting from ptp_{t} using the forward KL divergence. Likewise, let pt+1r(x)p^{r}_{t+1}(\mathbf{x}) be the distribution obtained using the reverse KL divergence, from ptp_{t}. Define Δtf\Delta^{f}_{t} and Δtr\Delta^{r}_{t} as the difference of log probability ratios across two categories x1\mathbf{x}_{1} and x2\mathbf{x}_{2}, obtained from the forward and reverse divergences respectively:

and Δtr\Delta_{t}^{r} is similarly defined. Then we have the following (for appropriate positive constants β\beta, δ1\delta_{1}, δ2)\delta_{2}):

Reverse KL modifies probability mass more aggressively than the forward KL. If x1\mathbf{x}_{1} and x2\mathbf{x}_{2} are such that, δ1≤pt(x1)=pt(x2)≤1−δ2\delta_{1}\leq p_{t}(\mathbf{x}_{1})=p_{t}(\mathbf{x}_{2})\leq 1-\delta_{2} (where δ1>0\delta_{1}>0, δ2>0\delta_{2}>0), but q(x1)≥q(x2)+βq(\mathbf{x}_{1})\geq q(\mathbf{x}_{2})+\beta, then, Δtr(x1,x2)>Δtf(x1,x2)\Delta^{r}_{t}(\mathbf{x}_{1},\mathbf{x}_{2})>\Delta^{f}_{t}(\mathbf{x}_{1},\mathbf{x}_{2}).

Reverse KL increases probability mass only on a subset of categories that equal target likelihoods. If x1\mathbf{x}_{1} and x2\mathbf{x}_{2} are such that, pt(x2)+β≤pt(x1)≤1−δ2p_{t}(\mathbf{x}_{2})+\beta\leq p_{t}(\mathbf{x}_{1})\leq 1-\delta_{2}, and q(x1)=q(x2)>c0⋅pt(x1)q(\mathbf{x}_{1})=q(\mathbf{x}_{2})>\mathbf{c}_{0}\cdot p_{t}(\mathbf{x}_{1}), where c0\mathbf{c}_{0} is a positive constant >1>1, then, Δtr(x1,x2)>Δtf(x1,x2)\Delta^{r}_{t}(\mathbf{x}_{1},\mathbf{x}_{2})>\Delta^{f}_{t}(\mathbf{x}_{1},\mathbf{x}_{2}).

Reverse KL aggressively reduces probability mass on less-likely categories in the target distribution. If x1\mathbf{x}_{1} and x2\mathbf{x}_{2} are such that, pt(x2)+β≤pt(x1)≤1−δ2p_{t}(\mathbf{x}_{2})+\beta\leq p_{t}(\mathbf{x}_{1})\leq 1-\delta_{2}, and q(x1)=q(x2)<c1⋅pt(x2)q(\mathbf{x}_{1})=q(\mathbf{x}_{2})<\mathbf{c}_{1}\cdot p_{t}(\mathbf{x}_{2}), where c1\mathbf{c}_{1} is a positive constant <1<1, then, Δtr(x1,x2)<Δtf(x1,x2)\Delta^{r}_{t}(\mathbf{x}_{1},\mathbf{x}_{2})<\Delta^{f}_{t}(\mathbf{x}_{1},\mathbf{x}_{2}).

A proof of Theorem 6.5 is shown in Section C.3. Essentially, this theorem enlists several cases where the forward KL modifies probability mass in different amounts across various categories, but the reverse KL acts disproportionately. In particular, case 1 says that the reverse KL exhibits more disproportionate probability mass changes on categories with equal likelihood pt(x)p_{t}(\mathbf{x}), due to the logarithmic dependency on the probability mass q(x)q(\mathbf{x}) (compared to the linear dependency for the forward KL). Case 2 says that when the target value q(x)q(\mathbf{x}) for two categories is much larger than the probability mass currently assigned to those categories, then the reverse KL can attempt to preferentially increase probability mass more in the category with a larger likelihood pt(x)p_{t}(\mathbf{x}) under certain conditions. Finally, case 3 shows that when the likelihood of a category is significantly larger than the target q(x)q(\mathbf{x}), the reverse KL is more effective at reducing this probability mass and re-distributing it to other categories within one update step. Finally, consider another special case, where the difference q(x)−pt(x)q(\mathbf{x})-p_{t}(\mathbf{x}) is identical for two categories x1\mathbf{x}_{1} and x2\mathbf{x}_{2}. In this case, while the forward KL will increase log probability ratios for both x1\mathbf{x}_{1} and x2\mathbf{x}_{2} equally, i.e., Δf(x1,x2)=0\Delta^{f}(\mathbf{x}_{1},\mathbf{x}_{2})=0, the reverse KL will prioritize the category with a higher pt(x)p_{t}(\mathbf{x}) value. These results highlight some scenarios under which the reverse KL can more efficiently re-organize probability mass across categories.

Discussion, Conclusion, and Limitations

We attempted to understand which components are particularly important for fine-tuning language models with preference data. Through extensive experiments on different fine-tuning problems in both didactic and LLM settings, we established that on-policy sampling is crucial for good performance especially when the peak in the ground-truth reward lies in less-likely regions of the reference policy initialization. That said, in practice, doing so requires preference datasets with broader coverage than the reference policy. We also showed that negative gradients can enable faster convergence and that objectives that induce a negative gradient are complementary to using on-policy sampling. Finally, we show that the notion of mode-seeking divergences unifies the notion of on-policy sampling and negative gradient. Our case study comparing forward and reverse KL divergences demonstrates the superiority of the reverse KL divergence in re-distributing probability mass efficiently, supporting our empirical findings pertaining to on-policy sampling and negative gradients.

While we conceptualize our observations, a limitation is that we don’t derive rigorous statistical guarantees in this work. As an example, we note that while the notion of concentrability coefficients (and associated guarantees) can potentially provide guarantees on on-policy sampling, to the best of our knowledge the notion of the negative gradient is not fully studied in the literature. We conjecture that negative gradient can perhaps be formalized statistically from the lens of providing a lower variance learning signal; it would be interesting for future work to formalize this. It would also be interesting to study more recent approaches based on minimax formulations (e.g., Munos et al. (2023); Yuan et al. (2024); Swamy et al. (2024); Chen et al. (2024)) in our empirical and conceptual framework. Next, while we consider the coverage of preference data relative to that of the reference policy in our study, this is a simplification that does not account for the coverage of the pre-training distribution which future work can incorporate. Finally, we remark that our study does not explore the effect of reward model quality, which tends to also play a central role in LLM fine-tuning. It would be interesting to extend our analysis to incorporate the role of reward model quality and parameterization.

Acknowledgements

We would like to thank Yi Su, Rishabh Agarwal, Zhang-Wei Hong, Young Geng, Abitha Thankaraj, Yuxiao Qu, So Yeon Min, Yutong He, Kevin Li, Sukjun Hwang, Khurram Yamin, Charlie Snell, Amrith Setlur, Kaylee Burns, Eric Mitchell, and others in CMU Russ Lab, CMU Auton Lab, Stanford IRIS Lab, and Stanford Ermon Group for discussions and feedback. AK thanks Aleksandra Faust, George Tucker, and Sergey Levine for informative discussions. This research is supported by computational resources from Google TPU Research Cloud (TRC) and the National Science Foundation. FT thanks Ruslan Salakhutdinov for insightful suggestions during this project. AS gratefully acknowledges the support of the NSF Graduate Research Fellowship Program.

References

Appendices

Our proposed framework also allows us to explain experiments and evaluations in several existing LLM fine-tuning results, and as a result, implies several practical guidelines for LLM practitioners. On the AlpacaFarm benchmark (Dubois et al., 2024), our results corroborate the gap between conditional supervised fine-tuning objectives such as binary FeedME and reward conditioning, and RL or contrastive training methods such as PPO and DPO: these results are perhaps even more extreme in that these conditional and weighted supervised fine-tuning objectives are not even able to outperform regular SFT. Methods that utilize on-policy sampling such as ReST (Gulcehre et al., 2023) and Quark (Lu et al., 2022) do outperform SFT but still underperform on-policy RL or on-policy contrastive training. The top-performing methods on the benchmark are offline DPO, which uses a negative gradient, and PPO, which leverages on-policy sampling.

Additionally, methods such as self-rewarding language models (Yuan et al., 2024), OAIF (Guo et al., 2024), DR-PO (Chang et al., 2024), Hybrid-DPO (Xiong et al., 2023), and RS-DPO (Khaki et al., 2024) couple on-policy sampling or rejection sampling with contrastive training objectives. These works corroborate our observation regarding the efficacy of on-policy sampling and negative gradients and how they are complementary. Approaches such as CRINGE (Adolphs et al., 2022) combine maximum likelihood with a token level contrastive loss term and show gains over solely utilizing supervised likelihood, corroborating our insights about negative gradients.

Concurrently to us, Xu et al. (2024) show that on many practical LLM fine-tuning problems offline DPO underperforms on-policy PPO. While we do not study the same LLM fine-tuning problems, the insights from this work corroborate our findings, which in turn extend insights from this work. For instance, this work observes that DPO can learn to find out-of-distribution responses, which is consistent with our analysis in Section 5.2.2 that offline DPO training might increase probability mass on the highly likely regions of πθ\pi_{\theta}, deviating significantly from the distribution of preferred responses p(yw∣x)p(\mathbf{y}_{w}|\mathbf{x}). To avoid this issue, this work prescribes an iterated DPO recipe where the reference policy (i.e., the SFT policy in their setting) is used to iteratively collect new samples for DPO training. Section 5.3 arrives at a similar conclusion that using on-policy samples for policy optimization, though we recommend collecting samples from the current policy and not the reference policy, which might fail to cover important regions of the space when the peak in the reward function appears farther away from the high-likely regions of the reference policy.

Appendix B Computational vs Wall-Clock Time Tradeoff for Various Methods

A natural takeaway extending the empirical results from Section 5.3 is that on-policy variants of contrastive approaches might provide for an better tradeoff between computation and wall-clock time. We perform a comparison of wall-clock time needed to run our experiments in Table 3. in particular, we found that on-policy DPO only requires 0.4 hours to converge, while offline DPO requires a wall-clock time of 1.3 hours to converge to the same solution in the Min Length scenario. In the Skew Length scenario, where the learned policy must deviate from the initial reference policy substantially, we find that while offline DPO can converge a bit quickly (0.12 hours), it flatlines at a sub-optimal solution (completion length of 11.8) as compared to on-policy DPO which takes merely 0.4 hours to reach a more optimal solution. This is far more time-efficient compared to other on-policy methods such as PPO and RWR that present a sampling bottleneck.

Appendix C More Details on Conceptual Unification and Theoretical Analysis

Here we provide proofs for the claims in 6.1. We will show that on-policy methods and offline constrastive methods, both are mode-seeking as opposed to supervised maximum likelihood approaches, which are mode-covering. This conceptually explains the differences in their behaviors that we observe in our experiments.

First, we prove Lemma 6.1, i.e., we want to show that on-policy RL methods and on-policy versions of weighted supervised learning methods optimize regularized version of a reverse KL-divergence.

Both on-policy RL algorithms and on-policy versions of weighted supervised learning, optimize the following loss function:

Following Appendix A.1 of Rafailov et al. (2023), there exists some policy π∗\pi^{*} such that we can express the reward function r(x,y)r(\mathbf{x},\mathbf{y}) as follows:

Note that Z(x)Z(\mathbf{x}) does not depend on πθ\pi_{\theta}. Therefore, minimizing LRL\mathcal{L}_{\text{RL}} with respect to πθ\pi_{\theta} is equivalent to optimizing the reverse KL-divergence. Since optimizing the reverse KL-divergence is mode-seeking, we see that on-policy RL algorithms have mode-seeking behavior. ∎

Next, we show that this is also the case for contrastive approaches as we prove Lemma 6.2.

First consider an input x\mathbf{x}. Consider the gradient update (with a small enough learning rate):

We shall prove that for all possible models θ\theta and for all tt, there always exists appropriate pairing of positive and negative samples (yw,yl)(\mathbf{y}_{w},\mathbf{y}_{l}), such that after taking the gradient update, we have:

The core idea behind this proof is the normalization of the probability simplex. We proceed with a combination of mathematical inducation and contradiction: assume that \omega_{t}\Big{|}_{{c_{2}>0}}\geq\omega_{t}\Big{|}_{{c_{2}=0}}, but for all possible pairings (yw,yl)(\mathbf{y}_{w},\mathbf{y}_{l}), we have \omega_{t+1}\Big{|}_{{c_{2}>0}}<\omega_{t+1}\Big{|}_{{c_{2}=0}}. We will show that this is not possible. To do this, we first derive the expressions for ωt+1\omega_{t+1} and then study under what conditions is it possible that for any pairing of positives and negatives, ωt+1\omega_{t+1} is smaller when c2>0c_{2}>0. The expression for ωt+1\omega_{t+1} is given by:

Now, define: f(t+1;\mathbf{y}_{w},\mathbf{y}_{l},\mathbf{x})=\omega_{t+1}\Big{|}_{{c_{2}>0}}-\omega_{t+1}\Big{|}_{{c_{2}=0}}, then we have:

Suppose that for all negatives yl\mathbf{y}_{l} for a given positive response yw\mathbf{y}_{w}, f(t+1;yw,yl,x)<0f(t+1;\mathbf{y}_{w},\mathbf{y}_{l},\mathbf{x})<0, then:

which is a contradiction since c1>0c_{1}>0. This means that there is at least one choice of yl\mathbf{y}_{l} for a given yw\mathbf{y}_{w}, for which Δ(yw,yl,x)≥0\Delta(\mathbf{y}_{w},\mathbf{y}_{l},\mathbf{x})\geq 0. This means that if f(t;yw,yl,x)>0f(t;\mathbf{y}_{w},\mathbf{y}_{l},\mathbf{x})>0 then f(t+1;yw,yl,x)>0f(t+1;\mathbf{y}_{w},\mathbf{y}_{l},\mathbf{x})>0. Averaging over x\mathbf{x} for all iterations then gives us the desired result, when starting from an initialization when starting from the same initialization for both the cases when c2>0c_{2}>0 and c2=0c_{2}=0.

Gradients for both DPO and IPO exhibit the form in Lemma 6.2. We now show that the gradient of both DPO and IPO takes the form shown in Equation 6.1. From Rafailov et al. (2023), the gradient of the DPO loss is:

Now we derive the gradient of the IPO loss. Define

Now we prove Lemma 6.3, which shows that supervised offline methods that optimize a maximum likelihood loss exhibit mode-covering behavior.

Offline supervised methods optimize the following loss function:

Hence offline supervised methods minimize the re-weighted forward KL-divergence. ∎

C.2 Characterization of Gradients of Forward and Reverse KL

The gradients of forward and reverse KL are given by:

We start with the definition of KL-divergence:

This proves Equation C.3. Similarly, we can write:

Now we calculate the partial derivative with respect to fjf_{j}:

Proof for Lemma 6.4. Now, if the logits ftf_{t} are being updated with gradient descent on loss L\mathcal{L}, the distribution at the next step pt+1p^{t+1} is given by:

Let’s consider what the characterization of pt+1p^{t+1} for the forward kl:

Noticing that the denominator is just a normalization constant, we can write this as:

Similarly the characterization of pt+1p^{t+1} for the reverse KL looks like:

C.3 Quantifying the Differences Between Forward and Reverse KL

We prove these statements case by case. First we prove the result for Case 1. In this scenario, we have the following:

The gap between Δf\Delta^{f} and Δr\Delta^{r} is now given by:

Now, we note by mean-value theorem, that there exists a c0∈[q(x2),q(x1)]c_{0}\in[q(\mathbf{x}_{2}),q(\mathbf{x}_{1})] such that,

Since dlog⁡p/dp=1/p>1d\log p/dp=1/p>1 for c0∈(0,1)c_{0}\in(0,1), we have that:

This quantity is positive when p(x1)>c0=δ1p(\mathbf{x}_{1})>c_{0}=\delta_{1}. This shows the result for Case 1.

Next we prove Case 2. In this setting we are given q(x1)=q(x2)≥p(x1)≥p(x2)+βq(\mathbf{x}_{1})=q(\mathbf{x}_{2})\geq p(\mathbf{x}_{1})\geq p(\mathbf{x}_{2})+\beta. In this case, the expressions for Δf\Delta^{f} and Δr\Delta^{r} are given by:

On the other hand, the expression for Δr(x1,x2)\Delta^{r}(\mathbf{x}_{1},\mathbf{x}_{2}) is given by:

Now we analyze each sub-term independently. First, we note the following expression for term (b):

where c′c^{\prime} is obtained by applying the mean value theorem on the difference log⁡p(x1)−log⁡p(x2)\log p(\mathbf{x}_{1})-\log p(\mathbf{x}_{2}). Now, since q(x1)≥c0⋅p(x1)q(\mathbf{x}_{1})\geq\mathbf{c}_{0}\cdot p(\mathbf{x}_{1}), log⁡q(x1)−log⁡p(x1)≥log⁡c0\log q(\mathbf{x}_{1})-\log p(\mathbf{x}_{1})\geq\log\mathbf{c}_{0}. Hence, if p(x2)p(\mathbf{x}_{2}) is upper bounded (i.e., when β\beta is large enough), then this difference (a)+(b)(a)+(b) in Equation C.8 is positive. Combining with Equation C.7, we note that: Δr(x1,x2)>0\Delta^{r}(\mathbf{x}_{1},\mathbf{x}_{2})>0, although Δf(x1,x2)<0\Delta^{f}(\mathbf{x}_{1},\mathbf{x}_{2})<0. This concludes the proof.

Next, we prove Case 3. Similar to the previous case, here Δf(x1,x2)=−η(p(x1)−p(x2))≤−ηβ<0\Delta^{f}(\mathbf{x}_{1},\mathbf{x}_{2})=-\eta(p(\mathbf{x}_{1})-p(\mathbf{x}_{2}))\leq-\eta\beta<0. In this case, expanding upon the expression of Δr(x1,x2)\Delta^{r}(\mathbf{x}_{1},\mathbf{x}_{2}) similarly as Case 2, in order to show the desired inequality Δr(x1,x2)<Δf(x1,x2)\Delta^{r}(\mathbf{x}_{1},\mathbf{x}_{2})<\Delta^{f}(\mathbf{x}_{1},\mathbf{x}_{2}), we need to prove that:

Then, to attain the desired inequality, we need:

Note that since c′′≥p(x2)c^{\prime\prime}\geq p(\mathbf{x}_{2}), as long as there exists a sufficiently small constant c1<1\mathbf{c}_{1}<1, such that:

the LHS of this equation will be smaller than the RHS α0\alpha_{0}. This proves the result for this case. ∎

Appendix D Additional Algorithmic Details

Online methods such as PPO or RWR that uses a learned reward model can suffer from gradient variance issues due to the differences in the reward score. In particular, adding or subtracting a baseline bb from the reward rϕ(x,y)r_{\phi}(\mathbf{x},\mathbf{y}) does not change the relative order of preferred or dispreferred responses; however, it can change the variance of the gradients, leading to instability of the optimization routine. To mitigate this, prior work (Ziegler et al., 2020) often normalizes the reward to have zero mean and unit variance. This can be done during the training process by computing the mean and variance of the reward from an online batch. Formally, let {x(i),y(i)}i=1B\{\mathbf{x}^{(i)},\mathbf{y}^{(i)}\}_{i=1}^{\mathcal{B}} be a batch of data with batch size B\mathcal{B} sampled from policy πθ\pi_{\theta}: one calculates the standardized reward rˉϕ(x(i),y(i))\bar{r}_{\phi}(\mathbf{x}^{(i)},\mathbf{y}^{(i)}) as:

where μ^=1B∑i=1Brϕ(x(i),y(i))\hat{\mu}=\frac{1}{\mathcal{B}}\sum_{i=1}^{\mathcal{B}}r_{\phi}(\mathbf{x}^{(i)},\mathbf{y}^{(i)}), σ^=1B−1∑i=1B(rϕ(x(i),y(i))2−μ^)2\hat{\sigma}=\sqrt{\frac{1}{\mathcal{B}-1}\sum_{i=1}^{\mathcal{B}}(r_{\phi}(\mathbf{x}^{(i)},\mathbf{y}^{(i)})^{2}-\hat{\mu})^{2}}.

D.2 IPO

IPO (Gheshlaghi Azar et al., 2023) is a contrastive algorithm similar to DPO. The key difference between them is their loss function: DPO optimizes the negative log-sigmoid loss whereas IPO optimizes an MSE-type objective. Formally, the IPO objective is:

Appendix E Method Hyperparameters

We did an extensive sweep over hyperparameters for individual offline and online algorithms for the language model experiments. We built our algorithm implementations off of the Huggingface TRL implementation (von Werra et al., 2020).

E.2 DPO (Rafailov et al., 2023)

E.3 Pref-FT (Dubois et al., 2024)

E.4 PPO (Schulman et al., 2017)

E.5 RWR

E.6 Iterated Best-of-N (Mukobi et al., 2023)

Appendix F Code For Running Experiments

We have made the code for this project public in this repository. The additional datasets used in our experiments are listed below:

We gratefully acknowledge the following codebases: TRL (von Werra et al., 2020), HALOs (Ethayarajh et al., 2023), minGPT (Karpathy, ), DrQ-v2 (Yarats et al., 2021a, b) and PAINT Xie et al. (2022).

Appendix G More on Didactic Bandit Problems

Here we present details of our didactic bandit problem. The reference policy shown in Figure 2 is obtained by collecting 10000 samples from a Cauchy distribution with location x0=−0.7x_{0}=-0.7, scale γ=0.4\gamma=0.4. Next, we clip this samples between the interval (−1,1)(-1,1), and divide the interval into 100 equally spaced bins. Starting from −1-1, we label these bins 0,…,990,\ldots,99 sequentially, and calculate the frequency of samples that fell into each bin. Finally, we define,

The reward functions R1\mathbf{R}_{1} and R2\mathbf{R}_{2} are defined as:

G.2 Algorithmic Details

In the bandit setting, we consider five algorithms: (1) Best-of-N, (2) IPO, (3) REINFORCE, (4) PPO and (5) RWR.

Best-of-N is similar to SuperHF (Mukobi et al., 2023)/ReST (Gulcehre et al., 2023) and in some way their simplification for the bandit setting. Best-of-N collects NN actions/responses for a prompt/state x\mathbf{x}, namely y1,y2,…,yN\mathbf{y}_{1},\mathbf{y}_{2},\ldots,\mathbf{y}_{N}. Next, we collect the rewards {R(x,yi)}i=1N\{\mathbf{R}(\mathbf{x},\mathbf{y}_{i})\}_{i=1}^{N}, and based on these rewards, choose the best action ybest=arg max⁡yiR(x,yi)\mathbf{y}_{\text{best}}=\operatorname*{arg\,max}_{\mathbf{y}_{i}}\mathbf{R}(\mathbf{x},\mathbf{y}_{i}). Finally, the loss function is the negative log-likelihood of this best action.

To show the efficacy of negative gradient, we can also directly add a term to this loss function minimizing log probability on dispreferred actions. Explicitly, we consider the following loss function:

where β\beta is a hyperparameter that we usually set to 1.01.0. We note that in practice this loss can quickly become unstable and proceed to −∞-\infty, in practice we only minimize the probability of dispreferred actions if it is above a certain threshold.

In contrast, IPO uses the loss function defined in Equation D.2. While regular IPO is an offline algorithm that uses a fixed preference dataset Dpref\mathcal{D}_{\text{pref}}, since we have access to the true reward function in the bandit setup, we create an online version of this algorithm as well. Here we also have a fixed set of prompts Dprompts\mathcal{D}_{\text{prompts}}, and given a policy π\pi, we can generate a preference dataset as follows: for each prompt x∈Dprompts\mathbf{x}\in\mathcal{D}_{\text{prompts}}, we can generate completions y1,y2,…,yN∼π(.∣x)\mathbf{y}_{1},\mathbf{y}_{2},\ldots,\mathbf{y}_{N}\sim\pi(.|\mathbf{x}). For any i≠ji\neq j, without loss of generality, assume R(x,yi)>R(x,yj)\mathbf{R}(\mathbf{x},\mathbf{y}_{i})>\mathbf{R}(\mathbf{x},\mathbf{y}_{j}). Then yi\mathbf{y}_{i} and yj\mathbf{y}_{j} are the preferred and dispreferred completions respectively, and we can form a preference dataset with all such (x,yw,yl)(\mathbf{x},\mathbf{y}_{w},\mathbf{y}_{l}) tuples.

For REINFORCE, we sample y∼πθ(.∣x)\mathbf{y}\sim\pi_{\theta}(.|\mathbf{x}), calculate the normalized reward R(x,y)‾\overline{\mathbf{R}(\mathbf{x},\mathbf{y})}, and use the following loss:

For PPO, let πgen\pi_{\text{gen}} be the policy used to generate the responses, and define r(x,y)=πθ(y∣x)πgen(y∣x)r(\mathbf{x},\mathbf{y})=\frac{\pi_{\theta}(\mathbf{y}|\mathbf{x})}{\pi_{\text{gen}}(\mathbf{y}|\mathbf{x})}. Then we use the following loss function:

where ϵ>0\epsilon>0 is a hyperparameter that controls how much we clip off-policy updates.

For RWR, we use the following loss function:

where β\beta is a hyperparameter, usually β=0.1\beta=0.1 in our experiments unless otherwise noted.

G.3 Experiment Details

For all experiments, we use a small GPT (Radford et al., 2018; Brown et al., 2020)-like transformer architecture (named ‘GPT-Nano’) with 0.9M parameters. We took the implementation from this public repository: minGPT (Karpathy, ).

Appendix H Additional Experiments on Synthetic LLM Setup

Figure 20 shows the performance of various algorithms in the mode length setup. We see that all algorithms perform similarly here.

H.2 Effect of On-policy Samples vs Samples from an Older Policy in Synthetic Length Settings

Figures 21 and 22 shows the effect of using on-policy samples vs samples from an older policy for RWR in the synthetic length experiments.

H.3 Sample Reuse in Synthetic LLM Settings

Figure 23 shows the effect of sample reuse in the Skew Length setting: similar to Min Length ( Figure 11), some sample reuse can improve sample efficiency. but excessive sample reuse can also hurt performance. Also, we see PPO with importance clipping is much better at sample reuse than Best-of-N.