Human Alignment of Large Language Models through Online Preference Optimisation

Daniele Calandriello, Daniel Guo, Remi Munos, Mark Rowland, Yunhao Tang, Bernardo Avila Pires, Pierre Harvey Richemond, Charline Le Lan, Michal Valko, Tianqi Liu, Rishabh Joshi, Zeyu Zheng, Bilal Piot

Introduction

Learning from feedback is a common approach to align the behaviour of artificial agents with human preferences (Knox & Stone, 2008; Griffith et al., 2013; Christiano et al., 2017; Warnell et al., 2018). In recent years, reinforcement learning from human feedback has become a common paradigm for fine-tuning large language models (Glaese et al., 2022; OpenAI, 2022).

The traditional approach to fine-tuning of large language models from human preferences is to learn a reward signal under the Bradley-Terry model (Bradley & Terry, 1952), and then perform reinforcement learning (RL) against this learnt reward signal (Christiano et al., 2017). Recently, Rafailov et al. (2023) proposed an alternative model-free approach, direct policy optimisation (DPO). DPO is mathematically equivalent to the method above, in the sense that the minimiser of the population loss is identical (Azar et al., 2023), yet DPO bypasses the learning of the reward signal. Both approaches, however, rely on the Bradley-Terry model.

Recently, two particular approaches to directly optimise against preference probabilities themselves, rather than a Bradley-Terry-derived reward function have been proposed. Identity preference optimisation (Azar et al., 2023, IPO) is an algorithm that aims to optimise preference probabilities against a fixed data distribution, and does so with an offline contrastive loss, as with DPO. By contrast, Nash-MD-PG (Munos et al., 2023) is an algorithm that aims to find a Nash equilibrium with respect to the preference probabilities, via online policy gradient updates against a regularised policy. Both algorithms have appealing properties, though on the face of it are unrelated: one is an offline contrastive algorithm optimising against a fixed policy, the other is an online algorithm aiming to find a Nash equilibrium.

In this work we bridge the gap between IPO and Nash-MD-PG, and use this theoretical bridge to propose a novel class of preference optimisation algorithms. Specifically, our principal contributions are: first, we identify several key factors of variation between IPO and Nash-MD-PG, including their use of offline/online data, contrastivity of their losses, and the nature of their equilibria. We use this understanding to identify the strengths of these algorithms, and combine these strengths into new preference optimisation algorithms. This allows us to propose Online IPO, an online variant of IPO. In addition, we establish a theoretical connection between Online IPO and self-play in the regularised two-player preference game used in deriving Nash-MD-PG. Second, motivated by Online IPO, we propose a preference optimisation algorithm, aiming to capture the best aspects of both IPO and Nash-MD-PG: IPO-MD, a version of IPO that interpolates between offline and online variants by using the lagged data distribution of Nash-MD-PG. Finally, we provide an experimental suite contrasting these algorithms in several applications, which provides detailed comparisons between the proposed methods and several baselines, with notable take-aways for practitioners.

Background

We begin by introducing the central preference optimisation problem, and relevant prior work.

We consider a (non-contextual) bandit problem (rather than a sequential setting), with finite action space Y\mathcal{Y}. This simplifies the notation considerably, and the ideas presented are straightforwardly extensible to the contextual/sequential setting where the actions/generations yy are conditioned on a state/prompt xx; we explain this in further detail in the descriptions of implementation details in the experiments.

Preferences. A preference function p:Y×Y→p:\mathcal{Y}\times\mathcal{Y}\rightarrow specifies pairwise preference probabilities between elements of Y\mathcal{Y}. Specifically, given y,y′∈Yy,y^{\prime}\in\mathcal{Y}, p(y≻y′)p(y\succ y^{\prime}) is the probability that yy is preferred to y′y^{\prime}. We will assume that preference functions satisfy the symmetry condition p(y′≻y)=1−p(y≻y′)p(y^{\prime}\succ y)=1-p(y\succ y^{\prime}).

At a high level, the goal of learning in this context is to find a policy in Δ(Y)\Delta(\mathcal{Y}) (the set of distributions over Y\mathcal{Y}) that tends to select actions y∈Yy\in\mathcal{Y} that are preferred over others. This is of course not a precise mathematical objective as stated, and there are several distinct ways in which this can be formalised, which we explore in greater detail below.

Data model. Actions are sampled from policies μ,μ′∈Δ(Y)\mu,\mu^{\prime}\in\Delta(\mathcal{Y}) as Y∼μ,Y′∼μ′Y\sim\mu,Y^{\prime}\sim\mu^{\prime} (we write Y,Y′∼μY,Y^{\prime}\sim\mu when μ=μ′\mu=\mu^{\prime}). Given a preference function pp, the preference distribution λp\lambda_{p} is defined for y,y′∈Yy,y^{\prime}\in\mathcal{Y} as the distribution corresponding to the following sampling procedure:

We will frequently make use of a collection of samples drawn from the data-generating policies and the preference distribution, which we denote by (yi+,yi−)i=1N(y^{+}_{i},y^{-}_{i})_{i=1}^{N}. In this case we will say the data is sampled from (μ,μ′,λp)(\mu,\mu^{\prime},\lambda_{p}) (or (μ,λp)(\mu,\lambda_{p}) when μ=μ′\mu=\mu^{\prime}). In offline settings, the data is typically generated from a fixed policy μ\mu, whereas in online settings, new data can be generated from a currently estimated policy π\pi.

2 RLHF with a Bradley-Terry reward model

and a reward function r^\hat{r} is learnt essentially by performing maximum likelihood in this model, given a collection of observed preferences (yi+,yi−)i=1N(y^{+}_{i},y^{-}_{i})_{i=1}^{N} sampled from (μ,λp)(\mu,\lambda_{p}), (approximately) maximising the objective

Policy optimisation then proceeds by aiming to maximise the expected reward r^\hat{r} under the π\pi, subject to a KL constraint against the initial policy:

3 Direct preference optimisation (DPO)

Rafailov et al. (2023) propose direct policy optimisation (DPO) as an alternative to RLHF as described above, noting that with the closed form in Equation (3), the learning of a reward function can be completely bypassed, instead reparametrising the optimal reward in terms of the optimal policy, substituting into Equation (1), and aiming to maximise the resulting objective with respect to π\pi:

The derivation of DPO implies that, mathematically, it yields the same optimal policy as the RLHF approach in Section 2.2, which is a regularised optimiser for the reward model r^\hat{r}.

4 Sequence Likelihood Calibration (SLiC)

Zhao et al. (2023) propose sequence likelihood calibration (SLiC) as an alternative to RLHF. Soon after, Liu et al. (2023) refine the SLiC loss by normalising the policy probabilities with the reference policy probabilities in order to get a regularised offline loss. In the remaining, we will refer to the following loss, with a dataset (yi+,yi−)i=1N(y^{+}_{i},y^{-}_{i})_{i=1}^{N} sampled from (μ,λp)(\mu,\lambda_{p}):

as the SLiC loss. This loss can be interpreted as an hinge-loss variation of DPO (Liu et al., 2023).

5 Identity policy optimisation (IPO)

Azar et al. (2023) note that in general, optimisation of the reward model described in Section 2.2, which is implicitly optimised by DPO, may not always yield an intuitively good policy for the preference probabilities pp. They also note that in practice, removal of the learnt reward function from the pipeline removes an important source of regularisation in the learning problem, and as such DPO may learn policies that are under-regularised, and converge to deterministic actions.

In order to circumvent these two issues, Azar et al. (2023) propose identity policy optimisation (IPO). The derivation begins from the objective of aiming to directly optimise preference probabilities (rather than a proxy reward) against a fixed policy μ\mu:

Similar to Equation (3), the optimal policy for this objective is expressible directly as

which Azar et al. (2023) use to derive the following equivalent offline IPO loss, with a dataset (yi+,yi−)i=1N(y^{+}_{i},y^{-}_{i})_{i=1}^{N} sampled from (μ,λp)(\mu,\lambda_{p}):

6 Nash-MD-PG

Rather than optimising preference probabilities against an offline dataset generated from some data-generating policy μ\mu, Munos et al. (2023) propose instead interpreting the objective in Equation (6) as one player’s objective in a two-player, constant-sum game. Specifically, two players select policies π1,π2∈Δ(Y)\pi_{1},\pi_{2}\in\Delta(\mathcal{Y}), with player ii receiving payoff

where π−i\pi_{-i} denotes the policy of the other player. Note that holding π−i\pi_{-i} fixed at μ\mu then yields an equivalent objective to Equation (6) for πi\pi_{i}. The proposal of Munos et al. (2023) is then to find a Nash equilibrium for this game, motivated by the idea that the policies in the resulting Nash equilibrium may be more robust, and are not overly specific to the data-generating distribution μ\mu. On the flip-side there may be some benefit to regularising the sampling distribution toward the data distribution, and Nash-MD-PG has a parameter β\beta that allows for this tradeoff, with self-play in one extreme (β=0)(\beta=0) and sampling from μ\mu in the other (β=1\beta=1). We denote this algorithm by Nash-MD-PG(β\beta).

The Nash-MD-PG(β\beta) algorithm, motivated by mirror-descent approaches to saddle-point computation, aims to do so by updating the policy π\pi in the direction of the following policy gradient:

Comparative discussion of preference optimisation algorithms

The algorithms DPO (Section 2.3), SLiC (Section 2.4), IPO (Section 2.5), and Nash-MD-PG (Section 2.6) are distinct along a number of axes.

Contrastivity. IPO is a contrastive algorithm, in that it labels a pair (Y,Y′)(Y,Y^{\prime}) according to the preference function (via λp\lambda_{p}) into a positive (preferred) and a negative (not preferred) example (Y+,Y−)(Y^{+},Y^{-}), and then updates the policy via gradients flowing through both samples, π(Y+)\pi(Y^{+}) and π(Y−)\pi(Y^{-}). By contrast, Nash-MD-PG is not contrastive; only the sampled action’s policy probability is directly updated based on the preference. Contrastive algorithms in general have the potential to be more data-efficient, making direct use of both samples in each policy update (see Appendix D for a condition under which a contrastive gradient estimate has lower variance than its non-contrastive counterpart).

Regularised Sampling. As discussed in Section 2.6, Nash-MD-PG allows for sampling from a mixture distribution between π\pi and the data-generating distribution μ\mu, and this can also lead to improved performance versus sampling from either policy.

It is clear that there are several combinations of the various properties of previous methods for which no algorithm yet exists, including combinations that could have advantages over previous work, for example an online contrastive method. Table 1 gives an overview of existing methods in terms of what we consider are strengths of these methods, and how the methods introduced in this paper fit in this context, combining these strengths.

Online IPO

We first aim to bridge the online/offline divide between IPO and Nash-MD-PG, by proposing a new variant of IPO, Online IPO, which makes use of an online, shifting data distribution.

To derive an update, we first start with the population loss for IPO, which is obtained by taking the minibatch loss in Equation (8), and taking an expectation over the dataset (under i.i.d. sampling from (μ,λp)(\mu,\lambda_{p})). This yields the (offline) IPO population loss

The Online IPO population loss is given by replacing the static data distribution, highlighted in red above, with the data distribution generated by the current policy, as displayed below

Here, SG[π]\texttt{SG}[\pi] denotes a stop-gradient around π\pi in the data distribution, meaning that although we generate data from π\pi to construct the loss, we do not differentiate through the data-generation process itself.

The population form of the Online IPO loss will be useful in the analysis that follows. We conclude our description of the approach by noting that the sample-based Online IPO loss coincides with the Offline IPOFor clarity, we will refer to the original formulation of IPO by Azar et al. (2023) as Offline IPO in what follows. loss in Equation (8), with the exception that the samples (yi+,yi−)i=1N(y^{+}_{i},y^{-}_{i})_{i=1}^{N} are drawn from the current policy π\pi.

2 Analysis

Before studying the performance of the newly derived Online IPO loss empirically, we pause to consider it from a theoretical perspective. In particular, we aim to understand for which policies this loss is stationary.

By the analysis for (Offline) IPO (Azar et al., 2023) summarised in Section 2.5, the gradient for the Online IPO loss in Equation (11) is zero iff π\pi satisfies

The minimiser of the online IPO objective is the Nash equilibrium of the regularised game described in Equation (9).

This is perhaps a surprising conclusion. Offline IPO is not motivated by game-theoretic considerations, yet by moving to an online variant, we have obtained a loss whose stationary point is precisely the Nash equilibrium of the preference game optimised by Nash-MD-PG. In fact, we can go further, and deduce a direct equivalence of expected updates between Online IPO, and self-play in this game. The proof for the following result is given in Appendix C.

The expected gradient of the Online IPO loss in Equation (11) is identical to the self-play update direction in the game with payoff as in Equation 9.

Here, self-play refers to the algorithm in which a policy is updated using gradient ascent on its expected payoff in the game described in Equation (9), against another player using the same policy:

Note that in expectation, this corresponds to Nash-MD-PG with β=0\beta=0; however, an important difference is that in Online IPO, updates are contrastive, which may result in variance reduction of the gradient estimate.

We have therefore established a close connection between Online IPO and Nash equilibria for the regularised game.

3 Online DPO

As the DPO and SLiC losses are similar to IPO, a natural question is whether online variants of DPO and SLiC are also related to the regularised game given Equation 9. We explore this question for online DPO in Appendix F. Lemma F.5 gives conditions for the stationary point of the regularised game and online IPO to be a stationary point of online DPO. However, the conditions seem difficult to satisfy: For example, Theorem F.6 says that there is no 2-action problem for which the condition is satisfied, except when preferences are uniform (p(1≻2)=12p(1\succ 2)=\frac{1}{2}). In this sense, apart from the trivial uniform-preference case, online DPO and online IPO are different objectives when ∣Y∣=2|\mathcal{Y}|=2.

As for stationary points of online DPO, we show that, under the Bradley-Terry model assumption, the RLHF solution (Equation 3) is a stationary point of online DPO (Theorem F.7), coinciding with that of offline DPO.

IPO-MD

Having established a connection between Online IPO and self-play, it is natural to consider whether we can improve on self-play by using regularised policies to generate the data, similar to how Nash-MD optimises preferences against a regularised adversary.

2 Analysis

By the analysis of Offline IPO (Azar et al., 2023), we have that any fixed point πβ∗\pi^{*}_{\beta} of IPO-MD(β\beta) must satisfy

Hence, the fixed point of IPO-MD(β)(\beta) coincides with the fixed point of Nash-MD-PG(β\beta).

Having established the equivalence of the stationary points of IPO-MD(β\beta) and Nash-MD-PG(β\beta), we now study their gradients more generally in the following result; see Appendix C for the proof.

The gradients of the algorithms Nash-MD-PG(β\beta) and IPO-MD(β\beta) are, respectively,

We recover the result mentioned earlier that when β=0\beta=0, IPO-MD(β=0\beta=0) (i.e., Online IPO) has a gradient aligned with that of Nash-MD-PG(β=0\beta=0) (i.e., Self-Play). Now as soon as β>0\beta>0, the gradients of these two algorithms are different. Interestingly however, as noted above, their fixed points remain the same.

We can also relate the fixed points themselves back to the original regularised game given in Equation (9), as described below.

Using the property in Equation (13), we have

which is the Nash equilibrium condition for the game described, as required. ∎

Experiments

We present our results on fine-tuning large language models where we compare our algorithms, online-IPO and IPO-MD, against recent baselines. In this section we only present the results for the online versions of IPO, DPO and SLIC to make the comparison against IPO-MD and Nash-MD-PG fair, and drop the corresponding ”online-” prefix for simplicity. We refer the reader to the appendix for results concerning the offline versions of those algorithms. Note that to aid interpretability and reproducibility, we now consider the contextual bandit case where the actions yy, also referred as generations, are conditioned on a prompt xx.

Setup and algorithms. We perform RLHF-style experiments where we initialise from a supervised-fine-tuned checkpoint, and then further fine-tune using one of the following algorithms: RL (regularised policy gradient), IPO, DPO , SLiC, Nash-MD and IPO-MD. These algorithms use either a learned reward model rϕr_{\phi} (RLHF) or a learned preference model pϕp_{\phi} (IPO, DPO, SLiC, Nash-MD, and IPO-MD). For our RL baseline, similarly to Munos et al. (2023), we use a regularised policy gradient update:

where rϕ(y∣x)r_{\phi}(y|x) is the reward model’s value for generation yy and context xx.

Implementation details. The contrastive offline algorithms such as IPO, DPO and SliC directly optimise the policy πθ\pi_{\theta} by minimising their respective losses over a pairwise dataset (xi,yi+,yi−)i=1N(x_{i},y_{i}^{+},y_{i}^{-})_{i=1}^{N} . In practice, we sample batches (xi,yi+,yi−)i=1B(x_{i},y_{i}^{+},y_{i}^{-})_{i=1}^{B} of size B≪NB\ll N and we minimise the following loss:

One important detail concerning IPO is that we use a simplified loss in our code. One can remark by expanding the square and removing terms that do not depend on θ\theta that the IPO loss is equivalent to:

The contrastive online algorithms such as IPO, DPO and SLiC use a trained preference model pϕp_{\phi}. To train pϕp_{\phi}, we use a pairwise dataset (xi,yi+,yi−)i=1N(x_{i},y_{i}^{+},y_{i}^{-})_{i=1}^{N} and follow the same protocol as Munos et al. (2023). Then, to train the policy πθ\pi_{\theta}, for each context xix_{i} of a batch, we sample two completely new generations (yi,yi′)∼πθ(y_{i},y_{i}^{\prime})\sim\pi_{\theta} according to πθ\pi_{\theta} and compute the preference pi=pϕ(yi≻yi′∣xi)p_{i}=p_{\phi}(y_{i}\succ y_{i}^{\prime}|x_{i}) via the preference model. Then, for each algorithm, we minimise the respective following loss:

Evaluation tasks and Models. In our experiments, we test all of the algorithms on an article summarisation task. We use the dataset described by Stiennon et al. (2020) that has been built from the TL;DR dataset (Völske et al., 2017). This is a dataset with pairwise preferences between alternate summaries. We train our preference and reward model on the train set DTrainD_{\texttt{Train}}, which contains 9282092820 examples. We evaluate reward and preference models on a test set of high confidence data DTestD_{\texttt{Test}} and use the checkpoints with the highest evaluation agreement score. To train the policies with online algorithms, we use prompts of the train set of the XSum dataset (Shashi et al., 2018).

We use T5X large language models (Roberts et al., 2022) to train our policies, rewards and preference models. The T5X models we use are auto-regressive transformers with an encoder-decoder architecture. All the details of the models architecture and the different sizes are provided in the documentation (Roberts et al., 2022). For the policy model, we use a large (L) encoder-decoder model (770M770M parameters). For the preference and reward models, we use an XL encoder-decoder model (3B3B parameters). To train reward and preference models we use the same losses and protocol as Munos et al. (2023). For summarisation, we initialise our policy with a T5X-L model and fine-tune it with supervised learning using the OpenAI dataset described by Stiennon et al. (2020). We call this supervised fine-tuned model the SFT. All our policies for summarisation are initialised with this SFT checkpoint.

Our evaluation pipeline is based upon the use of PaLM2 (Anil et al., 2023) as a judge for side-by-side comparisons. We sample responses for each of the policies trained by each algorithm from a test set of prompts, and ask PaLM2 to pick which one is better. We use validation and test prompts from the XSum dataset (Shashi et al., 2018) for evaluation for the summarisation task, which is the same procedure used by Munos et al. (2023). The evaluation prompt we use for the side-by-side comparison is:

You are an expert summary rater. Given a piece of text and two of its possible summaries, output 1 or 2 to indicate which summary is better.

Text - , Summary 1 - , Summary 2 -

We use cloud Tensor Processing Units (TPUs; Jouppi et al., 2023) in their version 5e5e for our hardware compute, either in configurations of 2×42\times 4 devices for training offline experiments, or 4×44\times 4 devices for online experiments. This setup typically yields speed of around 0.250.25 training steps per second (2424 hours per 20,00020,000 steps). We run our experiments with default parameters 10−410^{-4} for the learning rate, and a default total of 30,00030,000 training steps, using a batch size of 3232. The τ\tau factor is held constant throughout training, and we do not employ any warmup steps. We use the AdaFactor (Shazeer & Stern, 2018) optimizer with decay set to 0.80.8.

In this section, we present side-by-side evaluation scores between the following online algorithms: RL, IPO, DPO, SLiC, IPO-MD and Nash-MD-PG. Table 2 presents the side by side scores for the summarisation task. The checkpoints we evaluate are the best checkpoints except for RL that we use as a baseline of comparison to find the best checkpoints. The RL checkpoint is fixed and was chosen following the protocol of (Munos et al., 2023) after sweeping over 66 values of τ\tau ({0.01,0.02,0.05,0.1,0.15,0.2}\{0.01,0.02,0.05,0.1,0.15,0.2\}) and comparing the performance against the SFT checkpoint after 10k10k learner steps. To find the best checkpoints for the other algorithms, we evaluate every checkpoint of each algorithm against the RL checkpoint (over 20002000 prompts sampled from a validation split) at different learning steps values (we checkpoint every 20002000 learner steps for a total of 30k30k learner steps), regularisation parameter τ\tau (we sweep over 55 values {0.1,0.5,1.0,5.0,10.0}\{0.1,0.5,1.0,5.0,10.0\}) and also β\beta for IPO-MD and Nash-MD-PG (we sweep over 22 values 0.1250.125 and 0.250.25) and we take the best checkpoint. After finding the best checkpoint for every algorithm (see App. B.3), we re-run each method for 3 different seeds using the best hyperparamters. We then perform 9 side-by-side evaluation (i.e., 3×33\times 3 1vs1 evaluations between each of the 3 seeds for each pair of methods) using 20002000 prompts from a different validation split for each comparison. We report mean and standard deviation across these 9 comparisons.

On the summarisation task, looking only at the mean the best algorithm is IPO as it beats all the other algorithms on a side-by-side comparison. However, once we take into consideration the standard deviation IPO and IPO-MD’s performance becomes statistically indistinguishable, with both algorithms consistently beating all the other algorithms. This shows that those algorithms are indeed robust and are closer to a Nash optimum than the other algorithms. Those results are limited to a summarisation task and more experiments should be conducted to validate these results on a general conversational agent. However, we do think that summarisation is a good test bed to showcase the quality of human alignment algorithms because it is a complex and high-in-demand task.

2 Ablations and Additional Results on Summarisation

Figure 1 sweeps the regularisation parameter for IPO and DPO. It’s interesting to see for small values of regularisation IPO and DPO behave very similarly, but for larger values the score for IPO decays much faster. This matches the findings in Azar et al. (2023) that show that IPO has a much stronger regularisation effect than DPO as τ\tau gets larger.

Learning steps curve

Figure 2 shows the performance against RL for IPO Online as it trains. We can see that as the regularisation gets stronger, more training time is required to reach the best performance.

Conclusion

In this paper, we have identified several factors of variation, such as contrastivity, online/offline and regularised-sampling, between two recently proposed algorithms for preference optimisation, IPO and Nash-MD-PG. In doing so, we have introduced two new algorithms, online-IPO and IPO-MD, that combine different strengths of these existing algorithms, namely the loss function of IPO with the online sampling and regularised data distribution of Nash-MD. Theoretical analysis reveals a surprising equivalence at the level of expected update between online-IPO and self-play in a regularised two-player preference optimisation game. This important property is not possessed by online-DPO. Finally, our empirical investigation on a summarisation task also reveals that IPO-MD and online-IPO are promising approaches to preference optimisation at scale as they are the most robust algorithms. At the moment, our work is restricted to model of size 770M770M on a single task, future works will consist to scale our approach to a full conversational agent using a larger model (100+100+ billions parameter).

Acknowledgements. We are grateful for the collaborative environment at Google DeepMind. We would like to thank Shantanu Thakoor, Will Dabney, Doina Precup, Mohammad Gheshlaghi Azar, Olivier Bachem, Sertan Girgin, Matt Hoffman, Nikola Momchev, Bobak Shahriari, and Piotr Stanczyk.

References

APPENDICES

Appendix A Related Work

RLHF. Reinforcement learning from human feedback, as introduced in Christiano et al. (2017) (see also Bai et al. (2022a); Ouyang et al. (2022)) and often based on proximal policy optimisation (Schulman et al., 2017), is a critical element of making large language models helpful and aligned with preferences of human operators. While in itself it typically does not result in improved benchmark performance (Touvron et al., 2023), RLHF is nonetheless key to satisfying human-mediated interactions such as dialogue (Nakano et al., 2021; Ouyang et al., 2022). The complexity of the RLHF procedure (Casper et al., 2023), which can also be accomplished by multiple reinforcement learning algorithms such as actor-critic (Mnih et al., 2016; Glaese et al., 2022), has led to searching for algorithmic alternatives (Dong et al., 2023; Yuan et al., 2023; Zhao et al., 2023).

Recent developments in policy optimisation. In the special case of an additional Bradley-Terry model (Bradley & Terry, 1952) assumption for the human reward model, reinforcement learning has been found redundant; this allows for casting the problem of RLHF as a supervised one (Rafailov et al., 2023). Recent developments have focused on scaling the performance of such direct policy optimisation (DPO) methods (Tunstall et al., 2023; Ivison et al., 2023), as well as generalising its mathematical formulation (Azar et al., 2023; Wang et al., 2023; Tang et al., 2024). One of the key issues with direct policy optimisation - and RLHF in general - resides in their propensity to game or hack rewards (Amodei et al., 2016; Skalse et al., 2022; Pan et al., 2022; Pang et al., 2022) and become overoptimised, or under-regularised (Gao et al., 2022; Singhal et al., 2023; Kirk et al., 2023), which can be mitigated e.g. by using ensembling techniques (Wortsman et al., 2022; Eisenstein et al., 2023; Coste et al., 2023; Ramé et al., 2024). Rather than a reinforcement versus supervised learning dichotomy, it is the distinction between online and offline (Jaques et al., 2019) methods that seems more relevant in practice, as an online policy’s generations might start deviating substantially from the original dataset, leading to distribution shifts (Zhuang & Hadfield-Menell, 2020; Shin et al., 2023). Finally, alignment can also result from the self-play form of a two-player game, and not just single-policy optimisation. This perspective, taken in Munos et al. (2023); Swamy et al. (2024) has the added benefit of encompassing both online and offline settings, enabling smooth interpolation between them via a hyperparameter. In a similar vein, DPO has been shown to be able and improve thanks to iterated successive rounds (Yuan et al., 2024), expanding on known machine-critic alignment methods such as reinforcement learning from AI feedback (Bai et al., 2022b; Lee et al., 2023).

Appendix B Additional Experimental Results

Figure 3 and Figure 4 show a sweep over the regularisation parameter for Online and Offline DPO and IPO vs. RLHF and SFT respectively. It is interesting to note that the online versions significantly outperform the offline versions. This is understandable as this setting favours tremendously online methods. Indeed, the starting point is an already fine-tuned policy on summarisation data. Therefore, for online methods the first checkpoint is already able to sample good summaries which make it very easy to obtain good rewards/preferences and from there optimise either the reward/preference model.

B.2 Mixing ratio curve

Here in Figure 5 we show a sweep over the mixing ratio β\beta for IPO-MD and how it affects its win-rate over the RL baseline in summarisation. We draw the curves for a different learning rate and different learning steps than the optimal checkpoint of IPO-MD to show that most of the time the mixing ratio still help improve the performance.

B.3 Best hyperparameters found

We report here the optimal τ\tau, learning rate lr and, where applicable, mixture ratio β\beta used to obtain each algorithm’s best performance in Table 2.

Appendix C Proofs

We calculate expressions for the update directions directly, assuming that the policy π\pi is parametrised via a vector ϕ\phi.

Online IPO update. The update direction for online IPO is given by the negative of the derivative:

We can simplify the above by first considering the terms with a factor of τ−1\tau^{-1}:

Now considering the terms not involving τ\tau:

Self-play update. Self-play leads to the update direction given by

Therefore the expected online IPO update direction is exactly the same as that of self-play. ∎

Using the anti-symmetry of the preference model, i.e., p(y≻y′)=1−p(y′≻y)p(y\succ y^{\prime})=1-p(y^{\prime}\succ y), and combining terms, we have that

Appendix D Comparison of the variance of contrastive versus non-contrastive gradient estimates

Define the following gradient estimates based on non-contrastive vs contrastive loss functions:

These two estimate resemble those of the algorithms Self-Play (as implemented by Nash-MD-PG(β=0\beta=0)), which uses a non-contrastive loss, and online-IPO (equivalent to IPO-MD(β=0\beta=0)), which uses a contrastive loss, and we have

We know from the previous result that these two estimates have the same expectation. However their variance may differ. Their respective variance depends on a non-trivial combinaison of the policy representation and the specifics of the preference model. We now state a sufficient condition under which the contrastive gradient estimate has lower variance than its non-contrastive counterpart.

If the policy representation and the preference model are such that we have

Defining the random variables X1(y,y′):=−∇log⁡π(y)f(y,y′)X_{1}(y,y^{\prime}):=-\nabla\log\pi(y)f(y,y^{\prime}) and X2(y,y′):=∇log⁡π(y′)f(y,y′)X_{2}(y,y^{\prime}):=\nabla\log\pi(y^{\prime})f(y,y^{\prime}), we have

Now, let us compare the variance of these estimates when yy and y′y^{\prime} are independently drawn from the same distribution π\pi. The variance of the contrastive estimate is

This variance is lower than \mboxVar(X1)\mbox{Var}(X_{1}) as soon as \mboxCov(X1,X2)\mbox{Cov}(X_{1},X_{2}) is negative (this is the principle of antithetic variates for variance reduction).

We deduce that a sufficient condition for the contrastive estimate to have a lower variance than that of the non-contrastive estimate is:

This condition is true as soon as Equation (14) is satisfied. ∎

Appendix E Tabular example

Appendix F Supplementary Theoretical Study of Online DPO

In this section, we explore whether online DPO is related to the regularised game given Equation 9, given that Proposition 4.1 shows a similar relationship for online IPO and the regularised game.

Our strategy is to derive the gradient of the offline DPO objective (Lemma F.3), and then inspect whether different sampling distributions μ\mu and candidate solutions π\pi satisfy the KKT conditions (Boyd & Vandenberghe, 2004) for the minimising the objective. We can also explore online DPO, by inspecting the KKT conditions when the sampling distribution matches the candidate solution (μ=π\mu=\pi). For example, in Lemma F.5 we present conditions for the solution of the regularised game (see Equation 12) to also be a stationary point of online DPO. Given that the minimiser of the online online IPO objective is the Nash equilibrium of the regularised game given in Equation 9 (Proposition 4.1), Lemma F.5 gives us a condition for the online DPO problem to be equivalent (solution-wise) to the online IPO problem. However, the condition seems difficult to satisfy: For example, there is no 2-action problem for which the condition is satisfied with the exception when preferences are uniform (p(1≻2)=12p(1\succ 2)=\frac{1}{2}). In this sense, apart from the trivial uniform-preference case, online DPO and online IPO are different objectives when ∣Y∣=2|\mathcal{Y}|=2.

Regaring stationary points of online DPO, we show that, under the Bradley-Terry model assumption, the RLHF solution (Equation 3) is a stationary point of online DPO (Theorem F.7). We note in passing (Remark F.9) that the RLHF solution is not the only solution of offline DPO when μ(y)=0\mu(y)=0 for some y∈Yy\in\mathcal{Y}. In fact, there are infinitely many, with arbitrarily small probabilities for yy such that μ(y)>0\mu(y)>0. So while we can say that stationary points of online DPO are also stationary points of offline DPO, a solution for the offline DPO problem with a given μ\mu may not be a solution for corresponding the online DPO problem, because of what happens for yy such that μ(y)=0\mu(y)=0.

We let Δ∘(Y)≐{p∈Δ(Y):p(y)>0,∀y∈Y}\Delta^{\circ}(\mathcal{Y})\doteq\left\{p\in\Delta(\mathcal{Y}):p(y)>0,\forall y\in\mathcal{Y}\right\} be the interior of the simplex.

Assume that ∣Y∣<∞|\mathcal{Y}|<\infty and that p(y≻y′)=1−p(y′≻y)p(y\succ y^{\prime})=1-p(y^{\prime}\succ y) for all y,y′∈Yy,y^{\prime}\in\mathcal{Y}. The offline DPO problem with sampling distribution μ\mu can be written as

Moreover, for π∈Δ∘(Y)\pi\in\Delta^{\circ}(\mathcal{Y}),

The offline DPO objective is obtained by taking the expectation of the DPO loss (Rafailov et al., 2023) with Y+,Y−∼(μ,λp)Y^{+},Y^{-}\sim(\mu,\lambda_{p}):

Since μy,y′=μy′,y\mu_{y,y^{\prime}}=\mu_{y^{\prime},y}, sy,y′=−sy′,ys_{y,y^{\prime}}=-s_{y^{\prime},y} and p(y≻y′)=1−p(y′≻y)p(y\succ y^{\prime})=1-p(y^{\prime}\succ y), we have

The Nash equilibrium π∗\pi^{*} of the regularised game given in Equation 9 is a stationary point of online DPO iff:

Let π∗\pi^{*} be the Nash equilibrium of the regularised game given in Equation 9, the fixed point in Equation 12:

π∗\pi^{*} will be a stationary point of online DPO iff π=π∗\pi=\pi^{*} is a solution of the offline DPO problem with μ=π∗\mu=\pi^{*}. Since π∗∈Δ∘(Y)\pi^{*}\in\Delta^{\circ}(\mathcal{Y}), we can use Lemma F.3, π∗\pi^{*} is a solution of the offline DPO problem iff

and the result follows by using the fact that τ>0\tau>0. ∎

No 2-action regularised game given in Equation 9 has a Nash equilibrium that satisfies Equation 19, except for the regularised games with p(y1≻y2)=12p(y_{1}\succ y_{2})=\frac{1}{2}.

Let Y={1,2}\mathcal{Y}=\{1,2\}. For this two-action problem, we can write the preference matrix as

where Pyy′=p(y≻y′)P_{yy^{\prime}}=p(y\succ y^{\prime}).

Let α\alpha be such that π∗=(α,1−α)⊤\pi^{*}=(\alpha,1-\alpha)^{\top}. Then

Then the difference between both sides of Equation 19 for y=1y=1 is:

Now, if we let ε≐12−p\varepsilon\doteq\frac{1}{2}-p, we can see that for p<12p<\frac{1}{2}

so (considering the analogous case for p<12p<\frac{1}{2})

Therefore, if p≠12p\neq\frac{1}{2}, we cannot satisfy Equation 19 for y=1y=1 and y=2y=2 (note that α=0\alpha=0 or α=1\alpha=1 satisfy the equation for only one yy). ∎

is a critical point for offline IPO for any μ\mu, and a stationary point of online DPO.

Let us consider the offline DPO problem with sampling distribution μ\mu. Assuming that the preferences admit a Bradley-Terry model, we know from Rafailov et al. (2023) that the solution of the offline DPO problem is given by

It follows that πr\pi^{r} is an offline DPO solution under any sampling distribution μ\mu, and if μ∈Δ∘(Y)\mu\in\Delta^{\circ}(\mathcal{Y}) then πr\pi^{r} is the only solution. In particular this means that πr\pi^{r} is also a stationary point for online DPO (by taking μ=πr\mu=\pi^{r}).

To see this, assume that there exists Y′⊂Y\mathcal{Y}^{\prime}\subset\mathcal{Y} such that Y′≠∅\mathcal{Y}^{\prime}\neq\emptyset and μ(y)=0\mu(y)=0 for all y∈Y′y\in\mathcal{Y}^{\prime}. For any α∈(0,1]\alpha\in(0,1], define πα∈Δ∘(Y)\pi^{\alpha}\in\Delta^{\circ}(\mathcal{Y}) by