Direct Language Model Alignment from Online AI Feedback

Shangmin Guo, Biao Zhang, Tianlin Liu, Tianqi Liu, Misha Khalman, Felipe Llinares, Alexandre Rame, Thomas Mesnard, Yao Zhao, Bilal Piot, Johan Ferret, Mathieu Blondel

Introduction

To maximise the benefits of large language models (LLMs) to society, it is important to align them with human expectations and values (Ouyang et al., 2022; Bai et al., 2022a; Bubeck et al., 2023). The first method introduced for alignment was reinforcement learning from human feedback (RLHF, Christiano et al., 2017; Stiennon et al., 2020), which trains a reward model (RM) from pairwise preferences and then optimises a policy against the RM via reinforcement learning (RL). More recently, direct alignment from preferences (DAP) methods have emerged as popular alternatives to RLHF, such as direct preference optimisation (DPO, Rafailov et al., 2023), sequence likelihood calibration with human feedback (SLiC, Zhao et al., 2023), and identity policy optimisation (IPO, Azar et al., 2023). In contrast to RLHF, the DAP methods directly update the language model (a.k.a. policy) {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{\bm{\theta}}} using pairwise preference data, making the alignment simpler, more efficient and more stable (Rafailov et al., 2023).

However, the preference datasets used in DAP methods are often collected ahead of training and the responses in the dataset are usually generated by different LLMs. Thus, the feedback in DAP methods is usually purely offline, as {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{\bm{\theta}}} cannot get feedback on its own generations over training. This is problematic because of the significant distribution shift between the policy that generated the dataset and the policy being aligned: we train on the distribution induced by {\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}{\rho}} but evaluate on the distribution induced by {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{\bm{\theta}}} in the end. In contrast, in RLHF, the RM provides online feedback to generations from {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{\bm{\theta}}} during the RL step. This practice leads to on-policy learning, which was shown to improve exploration and overall performance (Lambert et al., 2022).

Inspired by RL from AI feedback (RLAIF) (Bai et al., 2022b; Lee et al., 2023), we hereby propose Online AI Feedback (OAIF) for DAP methods. Our method inherits both the practical advantages of DAP methods and the online nature of RLHF. Specifically, when aligning an LLM policy {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{\bm{\theta}}}, we follow a three-step procedure: 1) we sample two responses to a prompt from the current policy {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{\bm{\theta}}}; 2) we obtain online feedback over the two responses by prompting an LLM to mimic human preference annotation; 3) we use this online feedback to update the model {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{\bm{\theta}}} through standard DAP losses. Our approach is depicted in Figure 1. Unlike methods proposed by Xu et al. (2023); Liu et al. (2023); Xiong et al. (2023), OAIF skips the RM training, and directly extracts the preference from an LLM.

To show the effectiveness of our proposal, we perform an extensive empirical comparison between OAIF, existing offline DAP methods and RLHF methods. Our experimental protocol uses both AI and human evaluation on standard LLM alignment tasks: TL;DR (Ziegler et al., 2019), Anthropic Helpfulness and Harmlessness (Bai et al., 2022a). To summarise, we make the following contributions.

We demonstrate the effectiveness and generality of OAIF for turning offline DAP methods (DPO, IPO, SLiC) into online methods. Our human evaluation shows that the average win rate of online DAP methods (DPO, IPO, SLiC) over offline versions of the same methods is ∼66%{\sim}66\%.

We confirm the usefulness of making DAP methods online: human raters favour DPO with OAIF (thus, online DPO) over SFT baseline, RLHF and RLAIF 58.00%58.00\% of time on the TL;DR task in 4-way comparisons.

We demonstrate the controllability of the LLM annotator, by injecting specific instructions into the prompts. We use response length as a test-bed. By asking the LLM annotator to prefer shorter responses, the average length of responses from the aligned policy is significantly shortened from ∼120{\sim}120 to ∼40{\sim}40, while its quality is still improved over the SFT baseline.

Background

LLM-based online feedback for DAP methods. The method we propose next, “Online AI Feedback” (OAIF), consists in using an LLM as an online annotator. Our method relies on the observation that LLMs can approximate well human labelling and can generate reliable preferences over responses (Lee et al., 2023). In recent concurrent work, Yuan et al. (2024) proposed a “self-rewarding” approach, in which the policy being aligned provides online feedback to itself. In comparison, OAIF can leverage feedback from any LLM, including ones stronger than the LLM being aligned. Swamy et al. (2024) also concurrently investigates the importance of online preference, but still relying on RMs.

In Table 1, we summarise the characteristics of OAIF and of the existing offline and online DAP methods.

Direct alignment from online AI feedback

Bridging the gap. As we saw, DAP methods are simple, do not require a separate RM, but they use preference data pre-collected offline. On the other hand, RLHF methods interact online with the language model being aligned, but they require policy gradient techniques to obtain an unbiased gradient estimate and a value function to reduce the variance. To bridge the gap between these two families of methods, we propose a simple yet effective way to make DAP methods online.

As pointed out by Ziegler et al. (2019), online data collection is crucial for aligning language models. To solve the aforementioned offline problem in DAP methods, we propose to collect preferences on-the-fly for responses generated by the language model being aligned. Naturally, using human feedback would be prohibitively expensive. Prior studies have shown that AI feedback is a reliable and effective approximation to human labellers, especially for pairwise preference labelling (Lee et al., 2023). We therefore propose to use an LLM as online annotator, in order to collect the preference over pairs of responses, sampled from {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{{\bm{\theta}}^{t}}} on-the-fly during its alignment. We refer to the proposed approach as OAIF, which stands for online AI feedback.

Annotating prompts with text-controllability. We adopt a pairwise prompting scheme to collect AI feedback, i.e. we instruct the LLM annotator to choose which response is preferred among a pair, as in Lee et al. (2023). To avoid position bias, we calculate scores for the two response possible orders and use the average as the final score. Since OAIF leverages prompting techniques to collect feedback, the reward signals or the preference function can be easily adapted by modifying the prompts (Sun et al., 2024). This offers high flexibility without incurring any extra computation (such as retraining the RM) compared to RLHF and RLAIF. For example, in our experiments, we show that we can control the response length by simply prompting the annotator to prefer shorter responses.

Experiments

To demonstrate the generality of OAIF, we experiment with three DAP methods: DPO, IPO and SLiC. Based on preliminary experiments, we set β=0.1\beta=0.1 in DPO, β=1.0\beta=1.0 in IPO, and β=0.002\beta=0.002 in SLiC. We sample responses with a temperature of 0.9 during training. We adopt Adafactor (Shazeer & Stern, 2018) as the optimiser, and set the batch size to 128 and the learning rate to 5⋅10−75\cdot 10^{-7}, with a warm-up period of 150150 steps for all experiments. We evaluate models by computing win rates, i.e. how often one model’s response is better than the other. For automatic evaluation, we apply the same prompting technique as above but with Gemini Pro (Gemini Team et al., 2023) to reduce the risk of over-fitting and reward hacking (Gao et al., 2023). The validity of Gemini Pro as the judge is explored in Appendix C. For human evaluation, three raters are presented with responses generated from a set of policy models. Each rater is then asked to independently score the responses’ quality (from 1 to 5 where 5 denotes the highest) and to pick the best one, and the average score is then used to compare the models.

2 How effective is OAIF for LLM alignment?

We start by examining the effectiveness of OAIF for DAP methods (that use online AI feedback), compared to their offline counterparts (that use pre-collected offline human preferences). As a sanity check, we track the win rate of DPO with OAIF (“Online DPO”) and vanilla DPO (“Offline DPO”) against the SFT baseline on TL;DR. The results are given in Figure 3, where the results for RLAIF and RLHF are provided as references.

Next, we evaluate OAIF on different tasks, i.e., TL;DR, Helpfulness and Harmlessness. We select the best performing online and offline DPO models according to both manual inspection and their development set win rate against the SFT baseline by Gemini Pro. We then report side-by-side human evaluations comparing online DPO and offline DPO in Table 2.

3 How does OAIF generalise to other DAP methods?

As shown in Algorithm 1, OAIF is compatible with arbitrary DAP loss functions. We therefore check the effectiveness of OAIF for IPO and SLiC. The side-by-side human evaluation results on TL;DR comparing the online and offline counterparts of these methods are given in Table 3.

Compared to their offline counterparts, DAP methods with OAIF achieve promising win rates, ranging from ∼64%{\sim}64\% to ∼71%{\sim}71\%. The consistent ineffectiveness of offline DAP methods confirms that the existence of the offline and off-policy issue in DAP methods and greatly hinders the performance of aligning LLMs. The consistent superiority of online DAP methods via OAIF against their offline counterparts demonstrates that OAIF is a general framework effectively addressing these challenges.

4 How do DAP methods using OAIF perform compared to RLHF/RLAIF?

Understanding the merits of DPO and RLHF is still a relatively open research question. We argue that comparing online DPO with RLAIF and RLHF, which is interesting on its own sake, can also contribute to answering this question.

We adopt similar experimental setups for RLAIF and RLHF as before, to make the comparison as fair as possible: we employ PaLM 2-L as the AI feedback model for RLAIF and use the same pre-collected preference dataset to train RMs for RLHF. Our training and optimisation procedures follow Lee et al. (2023). Figure 4(a) shows the human evaluation results, where online DPO is more preferred than the other methods, in 58%58\% of the time.

We emphasise that the RM used in RLAIF and RLHF is often not updated during policy training. As a result, its response assessment ability may not generalise, as the output distribution from {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{{\bm{\theta}}^{t}}} evolves. To verify this hypothesis, we also trained an online DPO with the same RM used for RLAIF. It outperforms RLAIF, but significantly underperforms online DPO with OAIF, with a win rate of <30%{<}30\% judged by Gemini Pro. This experimental result supports the superiority of using LLMs over RMs to provide online feedback. Synchronously retraining the RM is feasible theoretically (Ziegler et al., 2019), but this would greatly complicate the training pipeline and increase training cost.

Despite the great performance of OAIF compared to various baselines, we found that OAIF tends to produce significantly longer responses. This may affect the LLM and human evaluation as both evaluators often prefer long generations, referred to as “length bias” by Singhal et al. (2023). To avoid the effect of such bias on analysing the performance of OAIF, we group the responses by their length, and plot the average quality score of each group. The results in Figure 4(b) show that online DPO with OAIF provides responses of higher quality than the other methods at fixed length, which further validates the effectiveness of OAIF.

5 How does the size of the LLM annotator affect performance?

Another important dimension arising during our experiment is the size of the annotating LLMs. Previous experiments are all based on PaLM 2 L for feedback collection. To examine the feasibility of feedback from smaller LLM annotators, we then replicate online DPO experiments on TL;DR but with feedback from PaLM 2-XS and PaLM 2-S instead. Figure 5 shows the comparison to SFT baseline, offline DPO, RLAIF, and RLHF models we used, as in the previous experiments.

The size of the LLM annotator clearly has a significant impact on OAIF. Generally, as size increases, online DPO obtains better performance. Compared to the initial SFT model, online DPO with OAIF performs significantly better regardless of AI labeller model sizes, suggesting that even OAIF from a small LLM annotator is helpful in improving the performance of alignment. In particular, OAIF with PaLM 2-XS (i.e. an LLM annotator of same-size) achieves comparable performance to RLHF, although the latter learns from human feedback. Further human evaluation confirms this observation: OAIF with PaLM 2-XS obtains an overall quality score of 3.41 out of 5, slightly better than RLHF (3.38) and comparable to offline DPO (3.46).

6 How prompt-controllable is OAIF?

While the necessity of LLM alignment has been widely recognised, what to align them with is still under debate, as human expectations vary greatly across regions and cultures, and may evolve over time. This indicates that the human preference annotation might change dramatically and frequently. In RLHF, such changes require re-annotating the preference dataset and re-training the RM, leading to high cost. In contrast, as OAIF is obtained through prompting the LLM annotator, its reward signal could be adjusted by simply modifying the prompts.

To examine this, we choose to explore the controllability of the length of responses by modifying the prompts to the LLM annotators. We take the online DPO model {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{\bm{\theta}}} trained to be as helpful as possible in Section 4.2 as the reference. We further train another two online DPO models with the same experiment setup, but in which the annotator is prompted to favor “helpful and short” and “helpful and very short” responses. The exact prompts given to the LLM annotators are provided in Table 6 and Table 8.

We display the average length of responses over training in Figure 6(a). The “short” and “very short” prompts given to the LLM annotator significantly shorten the responses from ∼120{\sim}120 tokens to ∼90{\sim}90 and ∼40{\sim}40 tokens respectively. This direct evidence demonstrates that the behaviour of policy {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{\bm{\theta}}} can be significantly changed through prompting the annotating LLM differently, and the degree of the changes can be controlled as well.

However, the above changes come at a cost. In Figure 6(b), we plot the win rate of the “helpful”, “helpful and short”, and “helpful and very short” models against the initial SFT baseline. We noticed that the shorter responses become much less helpful, as judged by Gemini Pro. Nevertheless, they still improve the performance of the aligned model over the SFT baseline. This finding is also confirmed by human evaluation: from “helpful”, “helpful and short” to “helpful and very short”, the average quality score drops from 4.08, 3.72 to 3.26, all outperforming the SFT baseline (3.19) still.

7 Can weaker AI labeller improve stronger LLM?

Section 4.5 shows that PaLM 2-XS could provide reasonable feedback that helps improving the alignment of LLMs, although it’s significantly smaller than PaLM 2-S/L. We argue that our approach offers an orthogonal solution to the weak-to-strong generalisation problem investigated by Burns et al. (2023). To verify that a weaker AI labeller can improve the performance of a stronger LLM model, we perform experiments using PaLM 2-S as the policy model (student) under two teacher settings: one with PaLM 2-XS (weaker teacher) and the other with PaLM 2-L (stronger teacher). The side-by-side automatic evaluation results on Helpfulness comparing against the SFT baseline and offline DPO are given in Figure 7. Our results suggest that OAIF from a weaker teacher indeed improved the alignment of PaLM 2-S, though they are less effective compared with the OAIF from a stronger teacher.

We hereby emphasise the essential difference between the setup investigated by Burns et al. (2023) and ours. In their work, the tasks for the teacher and student model are both supervised learning tasks, thus they are of equal difficulty. However, in our work, the role of teacher is a simpler discriminative task (labelling preference), whereas the student model being aligned is given a more difficult one (generating proper responses). Following this perspective, our method is actually closer in spirit to the generative adversarial network proposed by Goodfellow et al. (2020), but doesn’t train a particular discriminator.

Discussion

Limitations. In this work, we study only the shift between distributions over responses, e.g. {\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}{\rho}}({\bm{y}}|{\bm{x}}) and {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{{\bm{\theta}}^{t}}}({\bm{y}}|{\bm{x}}). However, the shifts also happen on the user prompt distribution pXp_{\mathcal{X}} and the ground-truth human value function. Although the prompt-controllability of OAIF raises a possible solution to later case, the shift of pXp_{\mathcal{X}} is still a challenge. Since we extract prompts from the given preference dataset, our study assumes an in-distribution of prompts used for evaluation, thus lacks of evaluating the performance of aligned LLMs on out-of-distribution prompts. In the meantime, the model aligned in Section 4 is always PaLM 2-XS, thus whether our conclusion holds after scaling up is not investigated. As pointed out by Bai et al. (2022a), it is harder to distinguish responses of higher quality. Therefore, how much can OAIF for responses from larger LLMs requires further study.

Self-annotating models. In all the experiments in Section 4, we aligned models {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{\bm{\theta}}} using preferences generated by a separate LLM annotator. Yet, technically speaking, the feedback could also be from the model {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{{\bm{\theta}}^{t}}} being trained at time-step tt. This method, used recently by Yuan et al. (2024), is promising as outputting responses and annotating preferences are two distinct tasks, the former being a generative task and the latter a discriminative task. However, one disadvantage of this approach is that the model architecture and size have to be the same. In contrast, the LLM annotator in OAIF can be of arbitrary nature: as shown in Section 4.5, an LLM annotator of larger size brings additional benefits. Therefore, we argue that the choice of LLM annotator should not necessarily be limited to the model being aligned, especially when an LLM annotator of larger size or higher quality is available.

Qualitative preference annotation from LLMs. While we used response length as a simple test-bed, the prompt-controllability of reward signals can be naturally extended to more qualitative desiderata. Human values (such as helpfulness and impartiality) are a typical example of qualitative desiderata. Moreover, one motivation for annotating preferences instead of quantitative scores by human labellers is indeed because grading how well a response follows human values is difficult. Our approach, however, shows that AI feedback can achieve the same goal by changing only the prompts to the LLM annotators. Our approach can be extended to align language models to other qualitative objectives without much input from human labellers.

Preference from real-time human feedback. In our work the online feedback is from LLM annotators, but it is technically plausible to replace them with real online users. In such case, the model can be aligned towards either a specific group of users or an individual user, and the key bottleneck becomes the sample efficiency for fine-tuning LLMs. During our experiment in Section 4.2, we found that the behaviour of a model can be visibly changed with ∼2,000{\sim}2,000 training steps, which requires ∼256,000{\sim}256,000 samples. To personalise an LLM, this amount of data is still way too much for an individual user to produce, which is a limitation of applying RLHF for single-user personalisation of LLMs. A common solution to improve sample efficiency is to use low-rank adaptation (LoRA) (Hu et al., 2021). However, aligning an LLM to a specific person requires several fundamental advances and we leave this to future research.

Conclusion

To circumvent the offline feedback problem in direct alignment from preference (DAP) methods, such as DPO, we proposed Online AI Feedback (OAIF), a simple and effective way to make DAP methods online via AI feedback. We carried out an extensive empirical evaluation, using both AI and human evaluation, which showed the effectiveness of DAP methods combined with OAIF, against their offline counterparts. We also exhibited the tendency of offline DAP methods to overfit, and in contrast the usefulness of OAIF as a way to mitigate reward overoptimization. We further verified the generality of OAIF, as our empirical results hold for three prominent DAP methods: DPO, IPO and SLiC.

Beyond the empirical evaluation of OAIF, our work also contributes the comparison of two types of methods: online DAP methods (e.g., online DPO) and RLAIF. Since the feedback comes from identical models in both learning algorithms, our experiment setup ensures that the AI feedback is of the same quality and that only the learning procedures differ. Our experimental results in various tasks show that online DPO outperforms RLAIF and RLHF, which further confirms the effectiveness of OAIF, compared to offline feedback. Moreover, we used response length as a test bed to demonstrate that the LLM annotator can be controlled easily using instruction prompts. This shows that OAIF can be used to achieve desirable alignment goals.

Overall, this work demonstrates the effectiveness and importance of OAIF for aligning LLMs, and paves the way for more scalable alignment strategies, requiring reduced human annotation effort.

Acknowledgement

We hereby acknowledge the enlightening discussion we had with Yao Fu for refining the initial design of our method, the invaluable assistance from Harrison Lee and Samrat Phatale on conducting experiments with RLAIF and RLHF, the insightful suggestions and feedback provided by Nino Vieillard which significantly contributed to enhancing the quality of our paper, as well as the dedication to developing the infrastructure essential for this project from Léonard Hussenot, Robert Dadashi, Geoffrey Cideron, Alexis Jacq, Sabela Ramos, Piotr Stanczyk, Sertan Girgin, Danila Sinopalnikov, Amélie Héliou, Nikola Momchev, Olivier Bachem, Sarah Perrin, Pier Giuseppe Sessa, Matt Hoffman, Bobak Shahriari.

Impact statements

We propose a new method to improve the alignment of AI with human values. Our method paves the way for more scalable alignment with reduced human efforts. Since we rely on AI feedback, to tackle other challenges in RLHF (Casper et al., 2023) and mitigate safety risks (Amodei et al., 2016), our approach must be considered within the larger context of responsible and safe AI.

Author contribution statement

Shangmin Guo: proposed the project idea, wrote the initial codebase, ran initial experiments, wrote prompts used in experiments, wrote the paper.

Biao Zhang: wrote the codebase, ran main experiments, further developed the prompts, wrote the paper.

Tianlin Liu: participated in discussions.

Tianqi Liu: contributed to the initial codebase, participated in discussions, gave comments on the paper.

Misha Khalman: performed human evaluation, participated in writing the experiment section.

Felipe Llinares: helped implement the initial codebase, helped setup the initial experiments.

Alexandre Ramé: contributed to the initial codebase, participated in discussions, gave comments on the paper.

Thomas Mesnard: helped implement initial codebase, gave comments on the paper.

Yao Zhao: contributed to the initial codebase, participated in discussions.

Bilal Piot: contributed to the codebase, participated in discussions, gave comments on the paper.

Johan Ferret, Mathieu Blondel: supervised the work, wrote the paper.

References

Appendix A Definition of On/offline and On/off-policy Learning in LLM Alignment

In this section, we are going to illustrate the online and offline, as well as the on-policy and off-policy aspects arising in DAP methods, RLHF, and RLAIF.

In RL, online learning, as opposed to offline learning, is about whether there are dynamic interactions between the policy and the environment (Levine et al., 2020):

Online RL refers to a scenario where the agent learns by directly interacting with the environment in real-time. Online RL is characterised by a continuous cycle of action, feedback, and learning, making it suitable for environments where the model can afford to learn through trial and error.

Offline RL, on the other hand, involves learning from a fixed dataset of experiences, without further interaction with the environment. This dataset comprises previous interactions, which may have been generated by the same agent or different policies.

Let’s now consider the setup of LLM alignment, following the notations we use in Section 2.

online if (yi+,yi−)=f(x,yi1,yi2)({\bm{y}}_{i}^{+},{\bm{y}}_{i}^{-})=f({\bm{x}},{\bm{y}}_{i}^{1},{\bm{y}}^{2}_{i}) where ff is an accessible preference function (either human labellers, RMs, or LLM annotators), and ({\bm{y}}_{i}^{1},{\bm{y}}^{2}_{i})\sim{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{{\bm{\theta}}^{t}}}(\cdot|{\bm{x}}_{i});

offline if yi+{\bm{y}}^{+}_{i} and yi−{\bm{y}}^{-}_{i} were generated from a potentially different policy {\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}{\rho}}, ahead of training.

Therefore, in RLHF and RLAIF, their RL step is consistently online, as y{\bm{y}} is sampled on-the-fly from the current policy, and the RM is always accessible to score y{\bm{y}} over training. We discuss the RM step in RLHF and RLAIF separately in Section A.3.

To sum up, online vs offline learning is about whether the responses are generated by the current policy and the feedback is given on-the-fly by a preference function , or the responses along with the feedback are pre-collected and kept fixed.

A.2 On-policy learning vs off-policy learning

The concepts of on-policy and off-policy learning in RL (Sutton & Barto, 2018) are given as follows:

On-policy learning refers to a scenario where the learning algorithm improves the policy based on data generated by the policy itself.

Off-policy learning, on the other hand, leverages data obtained from a different policy than the one being trained. Off-policy learning makes it possible to leverage the data generated by other models, or by previous versions of the policy.

On-policy if ({\bm{y}}_{i}^{+},{\bm{y}}^{-}_{i})\sim{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{{\bm{\theta}}^{t}}}(\cdot|{\bm{x}}_{i}), i.e. both yi+{\bm{y}}^{+}_{i} and yi−{\bm{y}}^{-}_{i} are sampled from {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{{\bm{\theta}}^{t}}} with xi{\bm{x}}_{i} as the input.

Therefore, DAP methods are off-policy if preference data comes from {\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}{\rho}}. Note that the conclusion is still true even if {\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}{\rho}}={\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{{\bm{\theta}}^{0}}}, since {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{\bm{\theta}}} keeps changing over training and {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{{\bm{\theta}}^{t}}}\neq{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{{\bm{\theta}}^{0}}} for t≠0t\neq 0. By contrast, the approach proposed in this work is an on-policy alternative, as responses are sampled from the current policy at each training step.

As can be seen from the above definitions and the ones in Section A.1, for DAP methods, offline DAP is also off-policy, as yi+{\bm{y}}^{+}_{i} and yi−{\bm{y}}^{-}_{i} are not sampled from the current policy. As a side note, it is technically possible for the online DAP to be off-policy, for instance if leveraging both online and offline data, but this practice is seldom used as of now.

Regarding the RL step in RLHF and RLAIF, as shown by the objective function in Equation 4 as well as the common practice in RLHF and RLAIF, the response to be scored by the RM is always from {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{{\bm{\theta}}^{t}}}:

Therefore, the RL step in RLHF is on-policy. Although the RL step can be technically off-policy, if partially or exclusively learning from samples from different policies, we note that such practice is not widespread at the time of writing.

To sum up, the on-policy and off-policy learning is about whether the distribution over responses yi+{\bm{y}}^{+}_{i} and yi−{\bm{y}}^{-}_{i} learned from is {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{{\bm{\theta}}^{t}}}(\cdot|{\bm{x}}_{i}).

A.3 Distribution shift between RM training and inference

in-distribution samples, if {\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}{\rho}}={\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{{\bm{\theta}}^{t}}}, i.e. if doing online data collection (Ziegler et al., 2019);

out-of-distribution (OOD) samples, if {\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}{\rho}}\neq{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{{\bm{\theta}}^{t}}}, which is the most common practice in RLHF.

Appendix B Distribution Shift in Preference Data Curation

Next, we inspect the log-probability of the preferred response y+{\bm{y}}^{+}, the less preferred response y−{\bm{y}}^{-} and the off-policy response yˉ\bar{{\bm{y}}} using GPT-2 Large, i.e. {\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}{\pi}_{\bm{\theta}}}. As shown in Figure 8, there is a clear margin between the log-probability of on-policy and off-policy responses, where GPT-2 Large assigns significantly lower probabilities to generations from PaLM 2-S. Thus, the results verify the existence of the distribution shift between the on-policy and off-policy preference data. Moreover, our experiments in Section 4.2 on comparing online and on-policy learning with offline and off-policy learning also indirectly shows the significance of solving this problem.

Appendix C Alignment Accuracy of Gemini Pro

Lee et al. (2023) showed that the judgement of PaLM 2-L correlates significantly with human, thus we adopted PaLM 2-L for online feedback collection during the training. To reduce the risk of over-fitting, we resort to Gemini Pro (Gemini Team et al., 2023) instead for automatic evaluation at the test phase. However, the quality of Gemini Pro’s judgement is not well studied yet.

In this section, we explore the correlation of Gemini Pro’s judgement with human’s judgement on the three datasets explored. Following Lee et al. (2023), we report alignment accuracy which measures the accuracy of LLM-labelled preferences with respect to human preferences.

Table 4 shows that Gemini Pro achieves an average alignment accuracy of 70.21%, which performs comparably to PaLM 2 L (70.72%). These results support our use of Gemini Pro for the judgement.

Appendix D Win Rate of Online DPO and Offline DPO against SFT over Training on TL;DR by PaLM 2 L

Appendix E Prompts for LLM Evaluation and AI Feedback Labelling

In this section, we list the prompts used for OAIF and the automatic evaluation. Each prompt follows a pairwise selection paradigm (Lee et al., 2023), which includes both responses apart from the input context and asks LLM to select the preferred one. In practice, we instruct LLM to produce a preference distribution by computing the softmax of the log-probabilities of generating the tokens “1” vs. “2”. We treat the probability as the preference score, based on which we provide online AI feedback and compute the win rate.

Lee et al. (2023) observed that the order of the two responses when instantiating the prompt has non-negligible impact on the selection, i.e. the so-called positional bias. To address this issue, we average the distribution over “{response1} vs. {response2}” and “{response2} vs. {response1}”.