Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Shusheng Xu, Wei Fu, Jiaxuan Gao, Wenjie Ye, Weilin Liu, Zhiyu Mei, Guangju Wang, Chao Yu, Yi Wu
Introduction
Large Language Models (LLMs) derive their extensive language patterns and knowledge through pre-training on substantial textual datasets (Brown et al., 2020; OpenAI, 2023; Touvron et al., 2023; Chowdhery et al., 2023; Anil et al., 2023). To leverage the formidable capabilities of LLMs in practical applications, a growing amount of research has underscored the importance of aligning these models with human preferences (Agrawal et al., 2023; Kadavath et al., 2022; Shi et al., 2023; Liang et al., 2021; Sheng et al., 2019). Various methods have been developed for fine-tuning LLMs, with popular approaches including Supervised Fine-Tuning (SFT) (Peng et al., 2023) and Reinforcement Learning from Human Feedback (RLHF) (Ziegler et al., 2019; Stiennon et al., 2020; Ouyang et al., 2022). Typically, fine-tuning involves two phases: SFT to establish a base model, followed by RLHF for enhanced performance. SFT involves imitating high-quality demonstration data, while RLHF refines LLMs through preference feedback.
Within RLHF, two prominent approaches are reward-based and reward-free methods. Reward-based methods, pioneered by OpenAI (Ouyang et al., 2022; Ziegler et al., 2019; Stiennon et al., 2020), construct a reward model using preference data and then employ actor-critic algorithms like Proximal Policy Optimization (PPO) to optimize the reward signal. In contrast, reward-free methods, including Direct Preference Optimization (DPO) (Rafailov et al., 2023), RRHF (Yuan et al., 2023), and PRO (Song et al., 2023), eliminate the explicit use of a reward function. DPO, a representative reward-free method, expresses the reward function in a logarithmic form of the policy and focuses solely on policy optimization.
Notably, the most successful applications like ChatGPT (OpenAI, 2022) and Claude (Antropic, 2023) are produced by the reward-based RLHF method PPO, while strong performances in academic benchmarks often result from the reward-free RLHF method DPO (Rafailov et al., 2023; MistralAI, 2023). This discrepancy raises two fundamental questions: 1) Is DPO truly superior to PPO in the RLHF domain? and 2) Can the performance of PPO be substantially improved in common RLHF benchmarks? In this paper, we delve into these questions. Through theoretical and empirical analysis, we uncover the fundamental limitations of DPO and explore critical factors that enhance the practical performance of PPO in RLHF.
First, our theoretical examination reveals that DPO might find biased solutions that exploit out-of-distribution responses. Empirically, we demonstrate that the performance of DPO is significantly affected by the distribution shift between the model outputs and the preference dataset. Second, we perform ablation studies on the algorithmic components of PPO and discover a collection of critical factors for PPO’s best RLHF performances, including advantage normalization, large batch size, and exponential moving average update for the reference model. Finally, we validate our findings through extensive experiments, including dialogue generation tasks and more challenging code generation tasks. These experiments feature diverse feedback types and difficulty levels. The results indicate that PPO consistently outperforms DPO across all experiments. Particularly, in the most challenging code competition tasks, PPO achieves state-of-the-art results. Specifically, on the CodeContest dataset (Li et al., 2022), our PPO model with 34B parameters outperforms AlphaCode-41B (Li et al., 2022), exhibiting a 10@1k improvement from 16.4% to 22.4%.
Related Work
Large language models (LLMs) trained on large datasets acquire surprising capabilities (Brown et al., 2020; OpenAI, 2023; Touvron et al., 2023; Chowdhery et al., 2023; Anil et al., 2023; Kaplan et al., 2020; Brown et al., 2020). To leverage these capabilities to real applications, pre-trained LLM is further fine-tuned on specific tasks (Radford et al., 2019; Chung et al., 2022; Tay et al., 2023). Through fine-tuning with popular approaches such as SFT and RLHF, LLMs demonstrate impressive performance on established benchmarks (Touvron et al., 2023; OpenAI, 2023), aligning further with human preferences and societal well-being (Russell & Norvig, 2020; Russell, 2022).
This paper concentrates on RLHF methods, which can be broadly categorized into reward-based and reward-free approaches. Reward-based methods entail training a reward model on preference data in an initial phase (Gao et al., 2023; Ziegler et al., 2019; Stiennon et al., 2020; Ouyang et al., 2022). Subsequently, this learned reward model is utilized to provide a reward signal for online Reinforcement Learning (RL) algorithms such as PPO (Schulman et al., 2017). There exist previous works that have studied these methods through hyper-parameter tuning and analyzed the effects of the quality reward model quality (Zheng et al., 2023; Casper et al., 2023). In contrast, reward-free methods offer a simpler training procedure by directly training LLMs on preference data or ranking data to distill human preference (Yuan et al., 2023; Liu et al., 2023; Touvron et al., 2023; Rafailov et al., 2023; Song et al., 2023; Dong et al., 2023; Hong et al., 2024). Among these reward-free methods, DPO (Rafailov et al., 2023) has demonstrated strong performances and become popular in the community (MistralAI, 2023; Chen et al., 2024; Yuan et al., 2024). Recent work discussed the performance gap of DPO and PPO on synthetic contextual bandits (Li et al., 2023). In this paper, We analyze the limitations of DPO theoretically and empirically, and explore the key factors for PPO training.
Concurrent efforts have been undertaken to avoid reward model overoptimization (Ramé et al., 2024), facilitate alignment data generation (Lee et al., 2023; Yang et al., 2023), and implement resource-efficient RLHF systems (Yao et al., 2023; Santacroce et al., 2023). These works complement our study and can be seamlessly integrated into our implementation. Previous works have explored the implementation details of PPO for LLMs (Zheng et al., 2023; Ramamurthy et al., 2023). Our paper extends its investigations with additional RLHF techniques, optimizing PPO performance to surpass its reward-free counterpart, DPO. Our work is also closely related to studies on algorithm implementation in the RL community (Engstrom et al., 2020; Andrychowicz et al., 2021; Yu et al., 2022). However, our findings provide further insights into fine-tuning LLMs with a model size of up to 34B parameters.
Preliminary
Language Model. We consider an LLM as a policy parameterized by . is designed to follow user instructions to generate a text response . We only consider single-round conversations to simplify notations. Given a prompt , the LLM will generate response in an auto-regressive manner:
where is the -th token in the response and is tokens in the response before .
SFT. As an initial phase of alignment, the pre-trained model is enforced to imitate high-quality demonstration data (dialogue, summarization, etc.), which is usually referred to as Supervised Fine-Tuning (SFT).
RLHF. To further align the SFT model with human preference, prior works (Ziegler et al., 2019; Ouyang et al., 2022) proposed the Reinforcement Learning from Human Feedback (RLHF) procedure, which maximizes the following objective,
In the rest of this section, we will introduce two representative algorithms to optimize Eq. 2: a reward-based approach, PPO, and a reward-free approach, DPO.
PPO. We can directly adopt standard reinforcement learning methods for Eq. 2. In this paper, we chose PPO as the training algorithm. When is unknown, a reward model is first learned from human-labeled data to approximate . A common practice is to collect a dataset of preference pairs . and are responses to and marked as “win” and “lose” by human respectively. The distribution of the preference dataset is assumed to follow the Bradley-Terry model (Bradley & Terry, 1952; Christiano et al., 2017), i.e., the probability of response is better than is given by
where is the sigmoid function. Given , is trained by minimizing the negative log-likelihood of Eq. 3:
After a reward model is obtained, is replaced with and could be explicitly optimized with online RL algorithms. We note that there exist cases when a ground-truth reward is available, and thus reward modeling becomes unnecessary (Zhang et al., 2020; Sellam et al., 2020; Ramamurthy et al., 2023). In these cases, the reward function can be directly incorporated into Eq. 2. While we acknowledge other actor-critic algorithms can also be feasible (Mnih et al., 2016; Haarnoja et al., 2018), we follow the mainstream work (Ziegler et al., 2019; Stiennon et al., 2020) and focus on PPO (Schulman et al., 2017) for our analysis in this paper.
DPO. Instead of learning a reward model, Direct Preference Optimization (DPO) (Rafailov et al., 2023) optimizes the policy over preference data. DPO derived the closed-form solution of Eq. 2, which reveals the relationship between the reward and the optimal language model :
where is a partition function that only depends on prompt . According to Eq. 5, if maximizes , the underlying reward can be derived with
We remark that although Rafailov et al. (2023) performs a single-round DPO over the preference dataset, some recent works also adapt DPO to an iterative variant with a learned reward model (Xiong et al., 2023; Yuan et al., 2024). We also investigate the performance of iterative DPO.
Understanding the Limitation of DPO
In this section, we demonstrate that DPO may not be superior to PPO. Firstly, we theoretically demonstrate issues with the DPO training objective. Secondly, we illustrate that DPO is more susceptible to out-of-distribution (OOD) data through a synthetic example. Lastly, through experiments on a real preference dataset, we validate that the performance of DPO can be improved by mitigating the distribution shift between the model outputs and the preference dataset.
It is well-known that PPO could exploit potential failures in the learned reward model to achieve high rewards without meeting the actual human preference, often manifested as erroneous (Lewis et al., 2017) or overly complex outputs (Singhal et al., 2023). We argue that, though DPO avoids reward modeling, DPO has a similar generalization issue. In the following theorem, we will show that any solution found by PPO also minimizes the DPO objective Eq. 7, and thus, any solution found by PPO that exploits the reward model can also be found by DPO. Furthermore, DPO might discover solutions exploiting out-of-distribution data, posing a risk of deviating excessively from the reference policy even when the reference policy aligns well with human preferences.
2 Empirical Validation in A Synthetic Scenario
We design a synthetic scenario to validate Theorem 4.1 in practice. We create discrete spaces of prompts and responses, both of size 8. The policy and reward model are modeled as MLPs, which take a one-hot vector as input and output a categorical distribution of overall responses. We manually enforce the optimal response to be diagonal indices. The preference dataset is randomly created under this constraint and only covers limited preference pairs for each input. The resulting policies of DPO and PPO are shown in Figure 1. We can see that in practice, DPO and the learned reward model can assign high values to the response out of the distribution of preference dataset, which are marked using circles. In the case of DPO, the final model may assign higher probabilities than the reference model to these responses, which is not desirable as performance improvement on OOD responses could not be guaranteed. For example, in the red circles, DPO increases the probability from 0.11 to 0.23. In contrast, though the reward model has a similar misspecification issue, PPO can alleviate the issue with explicit KL regularization w.r.t. the reference model.
Practical Remark: From the analysis in this section, we attempt to provide insights to understand the performance of DPO in practice — DPO is prone to generating a biased policy that favors out-of-distribution responses, leading to unpredictable behaviors. We will further validate these insights through an experimental study involving LLMs on real preference datasets.
3 Experiments on Real Preference Datasets
In this section, we conduct experiments on real preference datasets and investigate two aspects that may influence DPO performance, including the base model and preference data used for DPO training.
Impact of The Base Model. When using SFT (Alpaca) as the base and reference model, we find that DPO performs poorly, producing only a safety rate and low helpfulness reward. We hypothesize that this is caused by the distribution shift between the training data of the base model, i.e., the Alpaca dataset, and the preference data, i.e., the SafeRLHF dataset. To study the impact, we further fine-tune SFT (Alpaca) on the SafeRLHF dataset with safe responses to obtain SFT (Safe). We then use SFT (Safe) as the reference model to re-train DPO from scratch. As shown in Table 2, resolving the distribution shift issue essentially increases the safety rate by and the helpfulness reward from to .
Sensitivity to Preference Data. There exist pairs in the SafeRLHF dataset where both and have the same safety label. After filtering out the dual-unsafe and dual-safe preference data in the dataset, the trained model could obtain a much higher safety rate. However, filtering the dual-safe preference data would largely hurt the performance of helpfulness. These results suggest that while DPO may derive advantages from eliminating noise or controversies in the training data, excessively discarding high-quality data could be detrimental to DPO performance.
Impact of Preference Data Distribution. While mitigating the distribution shift can be done with additional SFT, we also investigate whether collecting additional data with the base model could bring benefit. Specifically, instead of using the existing preference data, we generate new responses with SFT (Safe) and use a learned reward model for preference labeling. We further repeat this process and iteratively set the reference model as the latest DPO model in the last iteration. We denote this method as DPO-Iter. Remarkably, DPO-Iter achieves a comparable safety rate with PPO. This experiment again demonstrates that DPO could be improved by mitigating the distribution shift. However, it also obtains a much lower helpfulness reward compared to PPO.
Practical Remark: The performance of DPO could be improved by mitigating the distribution shift between the model and the preference dataset. To alleviate the issue of distribution shift and noisy data, we suggest adopting the iterative DPO method. One should carefully annotate the model-generated samples each time and then proceed to the next round of training. However, we will demonstrate in Sec. 6 that even with a nearly perfect annotator, the performance of DPO remains unsatisfactory in challenging tasks such as code generation.
Key Factors to PPO for RLHF
In this section, we investigate the key factors to the RLHF performance of PPO. We find three key techniques: (1) advantage normalization (Raffin et al., 2021), (2) large-batch-size training (Yu et al., 2022), and (3) updating the parameters of the reference model with exponential moving average (Ouyang et al., 2022). The first two techniques are widely adopted by the RL community but are not well-studied in the field of RLHF. The third is a technique that has received limited discussion in the literature, involving the gradual update of the reference model through an exponential moving average (Ouyang et al., 2022). This particular approach has the potential to yield additional performance enhancements.
Our PPO implementation is based on DeepSpeed-Chat (Yao et al., 2023), except that (1) we use a scalar reward for each response instead of dense rewards assigned on each token and (2) we omit the auxiliary SFT loss during PPO training because of the limited amount of data. This implementation includes common PPO techniques such as value loss clip and generalized advantage estimation (GAE) (Schulman et al., 2016). We list experiment details in Appendix A.2.
Experimental Setup. Our ablation experiments for PPO are carried out on a dialogue task HH-RLHF (Bai et al., 2022) as well as two code generation tasks: APPS (Hendrycks et al., 2021) and CodeContest (Li et al., 2022). HH-RLHF is a preference dataset in the form defined in Section 3 that aims to train a helpful and harmless LLM. APPS and CodeContest datasets are competitive programming datasets. Given a problem, the LLM should output a piece of executable code to solve this problem. The correctness is verified by test cases in the dataset, which can then generate reward signals or preference pairs for PPO and DPO training, respectively. We remark that these two types of tasks feature different types of reward signals: preference and direct reward feedback. The complete experimental setup is listed in Section 6. In the experiment result, we denote advantage normalization as Adv. Norm., large batch-size training as LargeBatch and exponential moving average of reference model update as Ref. EMA.
Analysis. The result of the ablation study is shown in Table 3. In Table 3, with a small batch size, baseline PPO improves over the SFT model on HH-RLHF and CodeContest dataset but shows significant performance degradation on the APPS dataset. Advantage normalization stabilizes PPO training and improves the performance of PPO. The most significant benefit is brought by using a large batch size, especially on code generation tasks. Lastly, using the exponential moving average for the reference model also brings additional benefits. The intuition behind this is that while the main LLM of PPO is rapidly changing, the reference model should also be updated accordingly. Otherwise, the learned model may be strongly regularized to be close to the SFT model, which can hurt performance in challenging tasks. Figure 2 further demonstrates that increasing the batch size of PPO consistently improves the performance across all difficulty levels in the APPS dataset. We also highlight that utilizing a small batch size, such as 64, in PPO training could negatively impact the performance of the base SFT model, resulting in a 33.7% performance level on the introductory scale. We remark that our findings are consistent with those developed in the RL community (Yu et al., 2022).
Benchmark Results
In this section, we conduct experimental validations to evaluate the performances of both DPO and PPO. Initially, our experiments focus on general dialogue tasks, specifically HH-RLHF and SafeRLHF. The primary goal is to improve the effectiveness of LLM by promoting constructive interactions and mitigating detrimental components within the model. Additionally, our investigation extends to demanding code generation tasks, namely APPS and CodeContest.
HH-RLHF (Bai et al., 2022) dataset consists of human preferences on AI assistant responses, encompassing 170k comparisons. In this dataset, we conduct experiments based on Llama2-7B. We evaluate the trained models using the OpenAssistant reward modelhttps://huggingface.co/OpenAssistant/oasst-rm-2-pythia-6.9b-epoch-1. Note that this model is only used for evaluation and is not involved during training. In addition, we adopt GPT-4 to compare the responses of different models. The prompt and evaluation details are listed in Appenidx B.
As shown in Table 4, except DPO and PPO, we also investigate other alignment methods such as RRHF (Yuan et al., 2023) and PRO (Song et al., 2023). The results demonstrate that PPO and DPO are much more preferred by GPT-4 than the chosen responses in the dataset and SFT model outputs, outperforming RRHF and PRO across all metrics. In this paper, we focus more on the performance of DPO and PPO. We observe that DPO-Iter performs better than DPO but worse than PPO. PPO consistently achieves a higher reward and higher win rates. We also use GPT-4 to compare the outputs of DPO and PPO directly, and the results are listed in Table 5, which demonstrates that GPT-4 prefers the responses of PPO.
SafeRLHF (Dai et al., 2023) dataset comprises over 30k entries of expert comparison data. Each entry in this dataset contains two responses to a question. In our experiments, we consolidate two preferences as mentioned in Section 4.3. For evaluation, we borrow the official reward model and cost modelhttps://github.com/PKU-Alignment/safe-rlhf, which are trained to evaluate helpfulness and harmlessness, respectively.
The results on SafeRLHF are listed in Tab 6. Experiments indicate that after alignment, both DPO and PPO can generate responses with less harm, while PPO’s responses are more helpful.
APPS (Hendrycks et al., 2021) is a description-to-code generation benchmark from competitive programming platforms. For each question, there are also test cases to verify the accuracy of generated codes. We use these test cases in the training set to provide feedback. For PPO training, the feedback could be directly used as a reward. We simply define the reward as 10 if the generated code passes all test cases. Otherwise, the reward is 0. For DPO, since there are no preference pairs, we adopt DPO-Iter. Specifically, we use the base model to sample 5 codes for each prompt and utilize the test cases to label the correctness of generated codes. It is worth noting that for many prompts, the base model may fail to sample any correct answer. In such cases, we use the correct solutions from the dataset as . We evaluate the results using pass@k, which is defined as the proportion of problems successfully solved by employing k generated programs for each problem.
As shown in Table 7. We conduct experiments on different model sizes. In particular, when using CodeLlama-34B as the base model, we achieved state-of-the-art results on the APPS dataset. We can observe that DPO-Iter fails to improve the SFT model performances on all the model sizes. In contrast, for PPO, as the model size increases, the improvement is more apparent. We remark that the feedback using test cases is nearly perfect. However, the performance of DPO-Iter remains unsatisfactory.
CodeContest (Li et al., 2022) is a more challenging competitive programming dataset consisting of several programming languages. Here, we only use Python code. We adopt a similar way to train PPO as in the APPS dataset. For DPO training, we construct the preference dataset by using the correct and incorrect codes provided by the dataset. To compare with previous work, we adopt k@n to evaluate the generated code, which means that n samples will be evaluated on public tests in the problem description, and k of them will be submitted for hidden tests.
The results are listed in Table 8. We obtained similar conclusions as in APPS. PPO improves the SFT model significantly, while DPO fails to generate any correct codes. After one epoch of training, the code written by the DPO model has achieved a pass rate of 0, we observe that the DPO model outputs many meaningless code snippets. The results also demonstrate that DPO-Iter performs worse compared to SFT. With the assistance of PPO, CodeLlama-34B has surpassed the previous state-of-the-art on this task, outperforming Alphacode with 41 billion parameters.
Conclusion
In this paper, we uncover the fundamental limitations of DPO and explore critical factors that enhance the practical performance of PPO in RLHF. Through theoretical and experimental analysis, we explore the limitations of DPO and find that DPO is sensitive to the distribution shift between the base model outputs and preference data. We suggest that iterative DPO is better than training on static data. However, we also find that DPO fails to improve the performance on challenging tasks such as code generation. Moreover, according to the ablation study, we summarize the key factors for PPO training, including advantage normalization, large batch size, and updating the parameters of the reference model with an exponential moving average. With our practical tuning guideline, PPO demonstrates robust effectiveness across diverse tasks and achieves state-of-the-art results in challenging code competition tasks.
There are also limitations in our work. The reward model is significant in the training processes of both PPO and DPO-Iter. However, in this paper, we have not delved into the discussion of how to effectively train a robust reward model. For the code competition task, we utilize the ground-truth reward for PPO training and the labeling of DPO-Iter. However, this does not affect the conclusions drawn in our paper, and we leave it as future works.
Impact Statements
Our study investigates a critical challenge in aligning Large Language Models (LLMs) with human preferences, emphasizing its societal impact, including the elimination of bias and the reduction of unfairness. The use of a public dataset ensures transparency, mitigating concerns related to privacy and ethical considerations. This research emphasizes our dedication to responsible AI practices, aiming to improve societal well-being by aligning LLMs with human values while upholding rigorous standards for privacy and ethics.
References
Appendix A Implementation Details
For DPO training, we use = 0.1 with a learning rate of 1e-6. We sweep the batch size and report the best performance. For HH-RLHF and SafeRLHF, we train DPO for two epochs. For code generation tasks, we train DPO for a single epoch, since it has led to a deterioration in performance.
A.2 PPO Details
During the PPO training phase, we separate the parameters of actor and critic, and set the learning rate to 1e-5 for the actor model and 5e-6 for the critic model. By default, we set the global batch size as 512, and 512 roll-out samples are split into 4 mini-batches to update the actor and critic models. We configure the sampling parameters to include a temperature of 1.0 and a top-k value of 200. The advantage estimation parameter in GAE and the RL discount factor are fixed at 1. We set the KL penalty coefficient as 0.1, with a clipping value of 20 for reward scores. We additionally adopt advantage normalization and value normalization to stabilize the training.
For HH-RLHF and SafeRLHF, we set the maximum generated tokens as 256 and adopted PPO training for 5 epochs. For APPS and CodeContest, we set the maximum generated tokens as 1024, and adopt PPO training for 16 epochs. The checkpoints with the highest reward/pass@k on the validation sets are selected.
Appendix B GPT-4 Evaluation
We adopt the same evaluation prompt with (Rafailov et al., 2023). The prompt is :
When using GPT-4 for evaluation, we randomly sampled 100 queries from the test set. And ask GPT-4 to compare the two responses. To minimize the impact of response position on comparison, we swapped the positions of the two responses and evaluated them separately. If the results of the two evaluations are inconsistent, we set the final result as a “Tie”.
Appendix C Additional Experiments
We conduct experiments to assess the impact of distribution shift by varying the reference model. The results are listed in Table 9 and Table 10. Llama2-7B-SFT(Safe) and Codellama13B-SFT are models that are closer to the preference dataset in the Safe-RLHF and APPS dataset, respectively. The results indicate that DPO is more affected by the distribution shift than PPO.
C.2 Varying β𝛽\beta
In Table 11, We explore the impact of on the HH-RLHF and APPS datasets. On the HH-RLHF dataset, we evaluate the model using the OpenAssistant reward metric. On the APPS dataset, we report the average pass@5 score. The results indicate that having too large may harm the performance of both DPO and PPO. A value of 0.1 consistently performs well across various models and tasks.
C.3 Varying Preference Dataset
We train the model on a subset of the HH-RLHF preference dataset. The results are shown in Table 12. The results suggest that the performance of both PPO and DPO may be affected by the extent of coverage in the preference dataset. When training on the helpful-base subset, the performance of DPO has dropped to be similar to that of the SFT model.
We also evaluate PPO on the HH-RLHF dataset by filtering dual-unsafe and dual-safe preference pairs. The results are listed in Table 13. We observe that PPO could also be affected by the composition of the preference dataset. Overall, PPO maintains a safe rate of over 92% cross all the settings, while DPO is more affected by the preference dataset.
When filtering dual-unsafe samples, the PPO model achieves significantly higher helpfulness rewards. We hypothesize that it is because the reward model can discern helpfulness at a more nuanced level. Upon further filtering of dual-safe samples, we observe that the model becomes conservative, often declining to respond to questions altogether. This phenomenon occurs because, after filtering both dual-unsafe and dual-safe samples, the reward model focuses solely on safety. And refusing to respond could always be a safe option.
C.4 Human Evaluation
We also include human evaluation to validate the preference-based tasks. The results are listed in Table 14. We ensure that each reference pairs are evaluated by 4 different persons. Human agree with GPT-4 evaluations at a rate of 60% and 61%, respectively. According to human evaluation results, PPO outperforms both DPO and DPO-Iter.