SimPO: Simple Preference Optimization with a Reference-Free Reward
Yu Meng, Mengzhou Xia, Danqi Chen
Introduction
Learning from human feedback is crucial in aligning large language models (LLMs) with human values and intentions , ensuring they are helpful, honest, and harmless . Reinforcement learning from human feedback (RLHF) is a popular method for fine-tuning language models to achieve effective alignment. While the classical RLHF approach has shown impressive results, it presents optimization challenges due to its multi-stage procedure, which involves training a reward model and then optimizing a policy model to maximize that reward .
Recently, researchers have been exploring simpler offline algorithms. Direct Preference Optimization (DPO) is one such approach. DPO reparameterizes the reward function in RLHF to directly learn a policy model from preference data, eliminating the need for an explicit reward model. It has gained widespread practical adoption due to its simplicity and stability. In DPO, the implicit reward is formulated using the log ratio of the likelihood of a response between the current policy model and the supervised fine-tuned (SFT) model. However, this reward formulation is not directly aligned with the metric used to guide generation, which is approximately the average log likelihood of a response generated by the policy model. We hypothesize that this discrepancy between training and inference may lead to suboptimal performance.
In this work, we propose SimPO, a simple yet effective offline preference optimization algorithm (Figure 1). The core of our algorithm aligns the reward function in the preference optimization objective with the generation metric. SimPO consists of two major components: (1) a length-normalized reward, calculated as the average log probability of all tokens in a response using the policy model, and (2) a target reward margin to ensure the reward difference between winning and losing responses exceeds this margin. In summary, SimPO has the following properties:
Simplicity: SimPO does not require a reference model, making it more lightweight and easier to implement compared to DPO and other reference-based methods.
Significant performance advantage: Despite its simplicity, SimPO significantly outperforms DPO and its latest variants (e.g., a recent reference-free objective ORPO ). The performance advantage is consistent across various training setups and extensive chat-based evaluations, including AlpacaEval 2 and the challenging Arena-Hard benchmark. It achieves up to a 6.4 point improvement on AlpacaEval 2 and a 7.5 point improvement on Arena-Hard compared to DPO (Figure 1).
Minimal length exploitation: SimPO does not significantly increase response length compared to the SFT or DPO models (Table 1), indicating minimal length exploitation .
Extensive analysis shows that SimPO utilizes preference data more effectively, leading to a more accurate likelihood ranking of winning and losing responses on a held-out validation set, which in turn translates to a better policy model. As shown in Table 1, our Gemma-2-9B-it-SimPO model achieves state-of-the-art performance, with a length-controlled win rate on AlpacaEval 2 and a win rate on Arena-Hard, establishing it as the strongest open-source model under 10B parameters. Most notably, when evaluated on Chatbot Arena with real user votes, our model significantly improved upon the initial Gemma-2-9B-it model, advancing from 36th to 25th place and ranking first among all 10B models on the leaderboard. As of September 16th, 2024.
SimPO: Simple Preference Optimization
In this section, we first introduce the background of DPO (§2.1). Then we identify the discrepancy between DPO’s reward and the likelihood metric used for generation, and propose an alternative reference-free reward formulation that mitigates this issue (§2.2). Finally, we derive the SimPO objective by incorporating a target reward margin term into the Bradley-Terry model (§2.3).
DPO is one of the most popular preference optimization methods. Instead of learning an explicit reward model , DPO reparameterizes the reward function using a closed-form expression with the optimal policy:
where is the policy model, is the reference policy, typically the supervised fine-tuned (SFT) model, and is the partition function. By incorporating this reward formulation into the Bradley-Terry (BT) ranking objective , , DPO expresses the probability of preference data with the policy model rather than the reward model, yielding the following objective:
where are preference pairs consisting of the prompt, the winning response, and the losing response from the preference dataset .
2 A Simple Reference-Free Reward Aligned with Generation
Using Eq. (1) as the implicit reward has the following drawbacks: (1) it requires a reference model during training, which incurs additional memory and computational costs; and (2) it creates a mismatch between the reward optimized in training and the log-likelihood optimized during inference, where no reference model is involved. This means that in DPO, for any triple , satisfying the reward ranking does not necessarily mean that the likelihood ranking is met (here is the average log-likelihood in Eq. (3)). In our experiments, we observed that only of the triples from the training set satisfy this condition when trained with DPO (Figure 4(b)). This observation aligns with a concurrent work , which finds that existing models trained with DPO exhibit random ranking accuracy in terms of average log-likelihood, even after extensive preference optimization.
Length-normalized reward formulation.
One solution is to use the summed token log probability as the reward, but this suffers from length bias–longer sequences tend to have lower log probabilities. Consequently, when is longer than , optimizing the summed log probability as a reward forces the model to artificially inflate probabilities for longer sequences to ensure receives a higher reward than . This overcompensation increases the risk of degeneration. To address this issue, we consider using the average log-likelihood as the implicit reward:
This metric is commonly used for ranking options in beam search and multiple-choice tasks within language models . Naturally, we consider replacing the reward formulation in DPO with in Eq. (3), so that it aligns with the likelihood metric that guides generation. This results in a length-normalized reward:
where is a constant that controls the scaling of the reward difference. We find that normalizing the reward with response lengths is crucial; removing the length normalization term from the reward formulation results in a bias toward generating longer but lower-quality sequences (see Section 4.4 for more details). Consequently, this reward formulation eliminates the need for a reference model, enhancing memory and computational efficiency compared to reference-dependent algorithms.
3 The SimPO Objective
Additionally, we introduce a target reward margin term, , to the Bradley-Terry objective to ensure that the reward for the winning response, , exceeds the reward for the losing response, , by at least :
The margin between two classes is known to influence the generalization capabilities of classifiers . This margin is termed home advantage in Bradley-Terry models . In standard training settings with random model initialization, increasing the target margin typically improves generalization. In preference optimization, the two classes are the winning and losing responses for a single input. In practice, we observe that generation quality initially improves with an increasing target margin but degrades when the margin becomes too large (§4.3). One of DPO’s variants, IPO , also formulates a target reward margin similar to SimPO. However, its full objective is not as effective as SimPO (§4.1).
Objective.
Finally, we obtain the SimPO objective by plugging Eq. (4) into Eq. (5):
In summary, SimPO employs an implicit reward formulation that directly aligns with the generation metric, eliminating the need for a reference model. Additionally, it introduces a target reward margin to help separating the winning and losing responses. In Appendix F, we provide a gradient analysis of SimPO and DPO to further understand the differences between the two methods.
Preventing catastrophic forgetting without KL regularization.
Although SimPO does not impose KL regularization, we find that a combination of practical factors ensures effective learning from preference data while maintaining generalization, leading to an empirically low KL divergence from the reference model. These factors are: (1) a small learning rate, (2) a preference dataset that covers diverse domains and tasks, and (3) the intrinsic robustness of LLMs to learn from new data without forgetting prior knowledge. We present KL divergence experiments in Section 4.4.
Experimental Setup
We perform preference optimization with two families of models, Llama-3-8B and Mistral-7B , under two setups: Base and Instruct. In this section, our goal is to understand the performance of SimPO vs. other preference optimization methods in different experimental setups. Our strongest model is based on Gemma-2-9B (Instruct setup) with a stronger reward model, RLHFlow/ArmoRM-Llama3-8B-v0.1 (Table 1). We will present and discuss these results in Appendix J.
For the Base setup, we follow the training pipeline of Zephyr . First, we train a base model (i.e., mistralai/Mistral-7B-v0.1, or meta-llama/Meta-Llama-3-8B) on the UltraChat-200k dataset to obtain an SFT model. Then, we perform preference optimization on the UltraFeedback dataset using the SFT model as the starting point. This setup provides a high level of transparency, as the SFT models are trained on open-source data.
For the Instruct setup, we use an off-the-shelf instruction-tuned model (i.e., meta-llama/Meta-Llama-3-8B-Instruct, or mistralai/Mistral-7B-Instruct-v0.2) as the SFT models. It is unclear whether the released instruct checkpoints have undergone supervised fine-tuning (SFT) or the complete RLHF pipeline. For simplicity, we refer to these checkpoints as SFT models. These models have undergone extensive instruction-tuning processes, making them more powerful and robust than the SFT models in the Base setup. However, they are also more opaque because their RLHF procedure is not publicly disclosed. To mitigate the distribution shift between SFT models and the preference optimization process, we generate the preference dataset using the SFT models following . This makes our Instruct setup closer to an on-policy setting. Specifically, we use prompts from the UltraFeedback dataset and regenerate the chosen and rejected response pairs with the SFT models. For each prompt , we generate 5 responses using the SFT model with a sampling temperature of 0.8. We then use llm-blender/PairRM to score the 5 responses, selecting the highest-scoring one as and the lowest-scoring one as . We only generated data in a single pass instead of iteratively as in . We also experimented with using a stronger reward model, RLHFlow/ArmoRM-Llama3-8B-v0.1 , to rank generated data, which yields significantly improved performance (see Appendix H and Appendix J). This is the reward model we used in our Gemma 2 experiments.
In summary, we have four setups: Llama-3-Base, Llama-3-Instruct, Mistral-Base, and Mistral-Instruct. We believe these configurations represent the state-of-the-art, placing our models among the top performers on various leaderboards. We encourage future research to adopt these settings for better and fairer comparisons of different algorithms. Additionally, we find that tuning hyperparameters is crucial for achieving optimal performance with all the offline preference optimization algorithms, including DPO and SimPO. Generally, for SimPO, setting between 2.0 and 2.5 and between 0.5 and 1.5 leads to good performance across all setups. For more details, please refer to Appendix B.
Evaluation benchmarks.
We primarily assess our models using three of the most popular open-ended instruction-following benchmarks: MT-Bench , AlpacaEval 2 , and Arena-Hard v0.1 . These benchmarks evaluate the models’ versatile conversational abilities across a diverse set of queries and have been widely adopted by the community (details in Table 2). AlpacaEval 2 consists of 805 questions from 5 datasets, and MT-Bench covers 8 categories with 80 questions. The most recently released Arena-Hard is an enhanced version of an MT-Bench, incorporating 500 well-defined technical problem-solving queries. We report scores following each benchmark’s evaluation protocol. For AlpacaEval 2, we report both the raw win rate (WR) and the length-controlled win rate (LC) . The LC metric is specifically designed to be robust against model verbosity. For Arena-Hard, we report the win rate (WR) against the baseline model. For MT-Bench, we report the average MT-Bench score with GPT-4 and GPT-4-Preview-1106 as the judge model. GPT-4-Preview-1106 produces more accurate reference answers and judgments compared to GPT-4. For decoding details, please refer to Appendix B. We also evaluate on downstream tasks from the Huggingface Open Leaderboard benchmarks , with additional details in in Appendix C.
Baselines.
We compare SimPO with other offline preference optimization methods listed in Table 3. Many recent studies have extensively compared DPO and PPO . We will leave the comparison of PPO and SimPO to future work. RRHF and SLiC-HF are ranking losses. RRHF uses length-normalized log-likelihood, similar to SimPO’s reward function, while SLiC-HF uses log-likelihood directly and includes an SFT objective. IPO is a theoretically grounded approach method that avoids DPO’s assumption that pairwise preferences can be replaced with pointwise rewards. CPO uses sequence likelihood as a reward and trains alongside an SFT objective. KTO learns from non-paired preference data. ORPO ORPO can directly train on preference data without the SFT stage. For fair comparisons, we start ORPO from the same SFT checkpoints as other baselines, which yields better results than starting from base checkpoints. introduces a reference-model-free odd ratio term to directly contrast winning and losing responses with the policy model and jointly trains with the SFT objective. R-DPO is a modified version of DPO that includes an additional regularization term to prevent exploitation of length. We thoroughly tune the hyperparameters for each baseline and report the best performance. We find that many variants of DPO do not empirically present an advantage over standard DPO. Further details can be found in Appendix B.
Experimental Results
In this section, we present main results of our experiments, highlighting the superior performance of SimPO on various benchmarks and ablation studies (§4.1). We provide an in-depth understanding of the following components: (1) length normalization (§4.2), (2) the margin term (§4.3), and (3) why SimPO outperforms DPO (§4.4). Unless otherwise specified, the ablation studies are conducted using the Mistral-Base setting.
As shown in Table 4, while all preference optimization algorithms enhance performance over the SFT model, SimPO, despite its simplicity, achieves the best overall performance across all benchmarks and settings. These consistent and significant improvements highlight the robustness and effectiveness of SimPO. Notably, SimPO outperforms the best baseline by 3.6 to 4.8 points on the AlpacaEval 2 LC win rate across various settings. On Arena-Hard, SimPO consistently achieves superior performance, though it is occasionally surpassed by CPO . We find that CPO generates responses that are, on average, 50% longer than those generated by SimPO (See Table 10). Arena-Hard might favor longer generations due to the absence of a length penalty in its evaluation.
Benchmark quality varies.
Although all three benchmarks are widely adopted, we find that MT-Bench exhibits poor separability across different methods. Minor differences between methods on MT-Bench may be attributed to randomness, likely due to the limited scale of its evaluation data and its single-instance scoring protocol. This finding aligns with observations reported in . In contrast, AlpacaEval 2 and Arena-Hard provide more meaningful distinctions between different methods. We observe that the win rate on Arena-Hard is significantly lower than on AlpacaEval 2, indicating that Arena-Hard is a more challenging benchmark. Although our models excel on benchmarks, these evaluations have limitations, including restricted query space and potential biases from model-based evaluations. Efforts like WildBench aim to expand these spaces, where SimPO models demonstrate competitive performance.
The Instruct setting introduces significant performance gains.
Across all benchmarks, we observe that the Instruct setting consistently outperforms the Base setting. This improvement is likely due to the higher quality of SFT models used for initialization and the generation of more high-quality preference data by these models.
Both key designs in SimPO are crucial.
In Table 5, we demonstrate results from ablating each key design of SimPO: (1) removing length normalization in Eq. (4) (i.e., w/o LN); (2) setting the target reward margin to be 0 in Eq. (6) (i.e., ). Removing the length normalization has the most negative impact on the results. Our examination reveals that this leads to the generation of long and repetitive patterns, substantially degrading the overall quality of the output (See Appendix E). Setting to 0 yields also leads to a performance degradation compared to SimPO, indicating that it is not the optimal target reward margin. In the following subsections, we conduct in-depth analyses to better understand both design choices.
2 Length Normalization (LN) Prevents Length Exploitation
The Bradley-Terry objective in Eq. (5) essentially aims to optimize the reward difference to exceed the target margin . We investigate the relationship between the learned reward differences and the length difference between the winning and losing responses from the training set of UltraFeedback. We measure the difference of reward (; Eq. (4)) using the SFT model, the SimPO model, and a model trained with SimPO but without length normalization. We present the results in Figure 2(a) and observe that SimPO with LN consistently achieves a positive reward margin for all response pairs, regardless of their length difference, and consistently improves the margin over the SFT model. In contrast, SimPO without LN results in a negative reward difference for preference pairs when the winning response is shorter than the losing response, indicating that the model learns poorly for these instances.
Figures 2(b) and 2(c) illustrate the average log likelihood ( in Eq. (3)) versus response length on a held-out set for models trained with SimPO and SimPO without LN. The model trained without LN exhibits a much stronger positive Spearman correlation between likelihood and response length compared to SimPO, indicating a tendency to exploit length bias and generate longer sequences (see Table 11). In contrast, SimPO results in a Spearman correlation coefficient similar to the SFT model (see Figure 6(a)).
3 The Impact of Target Reward Margin in SimPO
We investigate how the target reward margin in SimPO affects the reward accuracy on a held-out set and win rate on AlpacaEval 2, presenting the results in Figure 3(a). Reward accuracy is measured as the percentage of preference pairs where the winning response ends up having a higher reward for the winning response than the losing response (i.e., ). We observe that reward accuracy increases with on both benchmarks, indicating that enforcing a larger target reward margin effectively improves reward accuracy. However, the win rate on AlpacaEval 2 first increases and then decreases with , suggesting that generation quality is not solely determined by the reward margin.
Impact of γ\gamma on the reward distribution.
We visualize the distribution of the learned reward margin and the reward of winning responses under varying values in Figure 2(b) and Figure 2(c). Notably, increasing tends to flatten both distributions and reduce the average log likelihood of winning sequences. This initially improves performance but can eventually lead to model degeneration. We hypothesize that there is a trade-off between accurately approximating the true reward distribution and maintaining a well-calibrated likelihood when setting the value. Further exploration of this balance is deferred to future work.
4 In-Depth Analysis of DPO vs. SimPO
In this section, we compare SimPO to DPO in terms of (1) likelihood-length correlation, (2) reward formulation, (3) reward accuracy, and (4) algorithm efficiency. We demonstrate that SimPO outperforms DPO in terms of reward accuracy and efficiency.
Although the DPO reward expression (with the partition function excluded) lacks an explicit term for length normalization, the logarithmic ratio between the policy model and the reference model can serve to implicitly counteract length bias. As shown in Table 6 and Figure 4(a), employing DPO reduces the Spearman correlation coefficient between average log likelihood and response length compared to the approach without any length normalization (referred to as ‘‘SimPO w/o LN’’). However, it still exhibits a stronger positive correlation when compared to SimPO. Note that this correlation does not fully reflect the generation length. Despite DPO showing a stronger correlation, the length of its generated responses is comparable to or even slightly shorter than those of the SimPO models. Please find more details in Appendix E.
DPO reward mismatches generation likelihood.
There is a divergence between DPO’s reward formulation, , and the average log likelihood metric, , which directly impacts generation. As shown in Figure 4(b), among the instances on the UltraFeedback training set where , almost half of the pairs have . In contrast, SimPO directly employs the average log likelihood (scaled by ) as the reward expression, thereby eliminating the discrepancy completely, as demonstrated in Figure 6(b).
DPO lags behind SimPO in terms of reward accuracy.
In Figure 4(c), we compare the reward accuracy of SimPO and DPO, assessing how well their final learned rewards align with preference labels on a held-out set. SimPO consistently achieves higher reward accuracy than DPO, suggesting that our reward design facilitates better generalization and leads to higher quality generations.
KL divergence of SimPO and DPO.
In Figure 5(a), we present the KL divergence between the policy model trained with DPO and SimPO and the reference model with different , measured on the winning responses from a held-out set during training. Figure 5(b) shows the corresponding AlpacaEval 2 LC win rate. Although SimPO does not apply any form of regularization against the reference model, the KL divergence of SimPO is reasonably small. Increasing reduces the KL divergence for both DPO and SimPO, with DPO exhibiting a more pronounced reduction at higher values. In this particular setting (Mistral-base), Figure 5(b) demonstrates that a smaller can improve AlpacaEval 2 performance, despite the higher KL divergence. We observe that in some settings (e.g., Llama-3-Instruct), a large (e.g., ) leads to better performance. We hypothesize that when the reference model is weak, strictly constraining the policy model to the reference model may not be beneficial. As a caveat, while we did not observe any training collapse or degeneration with proper tuning, in principle, SimPO could potentially lead to reward hacking without explicit regularization against the reference model. In such a scenario, the model might achieve a low loss but degenerate.
SimPO is more memory and compute-efficient than DPO.
Another benefit of SimPO is its efficiency as it does not use a reference model. Figure 5(c) illustrates the overall run time and per-GPU peak memory usage of SimPO and DPO in the Llama-3-Base setting using 8×H100 GPUs. Compared to a vanilla DPO implementation, DPO can be as memory efficient as SimPO if it were implemented to separate the forward passes of the reference model from the actual preference optimization. However, this implementation is not standard practice. SimPO cuts run time by roughly 20% and reduces GPU memory usage by about 10%, thanks to eliminating forward passes with the reference model.
Related Work
RLHF is a technique that aligns large language models with human preferences and values . The classical RLHF pipeline typically comprises three phases: supervised fine-tuning , reward model training , and policy optimization . Proximal Policy Optimization (PPO) is a widely used algorithm in the third stage of RLHF. The RLHF framework is also widely applied to various applications, such as mitigating toxicity , ensuring safety , enhancing helpfulness , searching and navigating the web , and improving model reasoning abilities . Recently, has highlighted challenges across the whole RLHF pipeline from preference data collection to model training. Further research has also demonstrated that RLHF can lead to biased outcomes, such as verbose outputs from the model .
Offline vs. iterative preference optimization.
Given that online preference optimization algorithms are complex and difficult to optimize , researchers have been exploring more efficient and simpler alternative offline algorithms. Direct Preference Optimization (DPO) is a notable example. However, the absence of an explicit reward model in DPO limits its ability to sample preference pairs from the optimal policy. To address this, researchers have explored augmenting preference data using a trained SFT policy or a refined SFT policy with rejection sampling , enabling the policy to learn from data generated by the optimal policy. Further studies have extended this approach to an iterative training setup, by continuously updating the reference model with the most recent policy model or generating new preference pairs at each iteration . In this work, we focus exclusively on offline settings, avoiding any iterative training processes.
Preference optimization objectives.
A variety of preference optimization objectives have been proposed besides DPO. Ranking objectives allow for comparisons among more than two instances . Another line of work explores simpler preference optimization objectives that do not rely on a reference model , similar to SimPO. proposes a method to jointly optimize instructions and responses, finding it effectively improves DPO. focuses on post-training extrapolation between the SFT and the aligned model to further enhance model performance. In this work, we compare SimPO to a series of offline algorithms, including RRHF , SLiC-HF , DPO , IPO , CPO , KTO , ORPO , and R-DPO , and find that SimPO can outperform them in both efficiency and performance. Recently, proposed a generalized preference optimization framework unifying different offline algorithms, and SimPO can be seen as a special case.
Conclusion
In this work, we propose SimPO, a simple and effective preference optimization algorithm that consistently outperforms existing approaches across various training setups. By aligning the reward function with the generation likelihood and introducing a target reward margin, SimPO eliminates the need for a reference model and achieves strong performance without exploiting the length bias. Extensive analysis demonstrates that the key designs in SimPO are crucial and validates the efficiency and effectiveness of SimPO. A detailed discussion of the limitations can be found in Appendix A.
Acknowledgments
The authors would like to thank Li Dong, Tianyu Gao, Tanya Goyal, Di Jin, Yuchen Lin, Kaifeng Lyu, Sadhika Malladi, Eric Mitchell, Lewis Tunstall, Haoxiang Wang, Wei Xiong, Zhen Xu, Libing Yang, Zhiyu Zhao, and members of the Princeton NLP group for their valuable feedback and discussions. We thank Niklas Muennighoff for his advice on training and reproducing training KTO models. We thank Haoran Xu for helping verify our CPO runs. Mengzhou Xia is supported by an Apple Scholars in AIML Fellowship. This research is also funded by the National Science Foundation (IIS-2211779) and a Sloan Research Fellowship.
References
Appendix A Limitations
More in-depth theoretical analysis. Despite the empirical success and intuitive motivation of SimPO, a more rigorous theoretical analysis is necessary to fully understand the factors contributing to its effectiveness. Additionally, we introduce an additional hyperparameter, the target reward margin, which requires manual tuning. Future work could explore how to determine the optimal margin automatically and provide a more theoretical understanding of SimPO.
Safety and honesty. SimPO is designed to optimize the generation quality of language models by pushing the margin between the average log likelihood of the winning response and the losing response to exceed a target reward margin. However, it does not explicitly consider safety and honesty aspects, which are crucial for real-world applications. Future work should explore integrating safety and honesty constraints into SimPO to ensure that the generated responses are not only high-quality but also safe and honest. The dataset used in this work, UltraFeedback , primarily focuses on helpfulness, and future research may consider a more comprehensive study utilizing larger-scale preference datasets and evaluation benchmarks that place a strong emphasis on safety aspects. Nonetheless, we observe that this method consistently achieves high TruthfulQA performance compared to other objectives in Table 9, suggesting its potential for safety alignment.
Performance drop on math. We observed that preference optimization algorithms generally decrease downstream task performance, particularly on reasoning-heavy tasks like GSM8k, as shown in Table 9. SimPO occasionally results in performance comparable to or worse than DPO. We hypothesize that this may be related to the choice of training datasets, hyperparameters used for training, or a mismatch of chat templates used for downstream task evaluations. One explanation is that the preference optimization objective may not be effectively increasing the likelihood of preferred sequences despite increasing the reward margin. first observed this phenomenon and point out that this can hinder learning from math preference pairs where changing one token can flip the label (e.g., changing to ). They propose a simple regularization strategy to add back a reference-model calibrated supervised fine-tuning loss to the preference optimization objective, and effectively mitigate this issue. Future work may consider integrating this regularization strategy into SimPO to improve performance on reasoning-heavy tasks.
Appendix B Implementation Details
We find that hyperparameter tuning is crucial for achieving optimal performance of preference optimization methods. However, the importance of careful hyperparameter tuning may have been underestimated in prior research, potentially leading to suboptimal baseline results. To ensure a fair comparison, we conduct thorough hyperparameter tuning for all methods compared in our experiments.
For the Base training setups, we train SFT models using the UltraChat-200k dataset with the following hyperparameters: a learning rate of 2e-5, a batch size of 128, a max sequence length of 2048, and a cosine learning rate schedule with 10% warmup steps for 1 epoch. All the models are trained with an Adam optimizer .
For the preference optimization stage, we conduct preliminary experiments to search for batch sizes in and training epochs in . We find that a batch size of 128 and a single training epoch generally yield the best results across all methods. Therefore, we fix these values for all preference optimization experiments. Additionally, we set the max sequence length to be 2048 and apply a cosine learning rate schedule with 10% warmup steps on the preference optimization dataset.
Method-specific training hyperparameters.
We have noticed that the optimal learning rate varies for different preference optimization methods and greatly influences the benchmark performance. Therefore, we individually search the learning rates in the range of [3e-7, 5e-7, 6e-7, 1e-6] for each method. Table 7 shows the detailed information on method-specific hyperparameters search ranges for baselines. There is a discrepancy between the KTO runs in their original paper, where the original runs use a RMSProp optimizer . We use an Adam optimizer for all the experiments. Table 8 shows SimPO’s hyperparameters used under each setting.
Decoding hyperparameters.
For AlpacaEval 2, we use a sampling decoding strategy to generate responses, with a temperature of 0.7 for the Mistral-Base setting following zephyr-7b-beta, https://github.com/tatsu-lab/alpaca_eval/blob/main/src/alpaca_eval/models_configs/zephyr-7b-beta/configs.yaml a temperature of 0.5 for the Mistral-Instruct setting following Snorkel-Mistral-PairRM-DPO, and a temperature of 0.9 for both Llama 3 settings. We grid search the temperature hyperparameter for the Llama-3-Base setting with DPO over , and fix it for all different methods. For Arena-Hard, we use the default greedy decoding for all settings and methods. For MT-Bench, we follow the official decoding configuration which defines different sampling temperatures for different categories.
Computation environment.
All the training experiments in this paper were conducted on 8×H100 GPUs based on the alignment-handbook repo. https://github.com/huggingface/alignment-handbook
Appendix C Downstream Task Evaluation
To examine how preference optimization methods affect downstream task performance, we evaluate models trained with different methods on various tasks listed on the Huggingface Open Leaderboard . These tasks include MMLU , ARC , HellaSwag , TruthfulQA , Winograd , and GSM8K . We follow the established evaluation protocols and present the results for all models in Table 9. Generally, we find that preference optimization’s effect varies across tasks.
Compared to the SFT checkpoint, we find that all preference optimization methods generally maintain MMLU performance with minimal decline. In this aspect, SimPO is largely comparable to DPO.
Reading comprehension and commonsense reasoning improves.
For ARC and HellaSwag, preference optimization methods generally improve performance compared to the SFT checkpoint. One hypothesis is that the preference optimization dataset contains similar prompts to these tasks, which helps the model better understand the context and improve reading comprehension and commonsense reasoning abilities.
Truthfulness improves.
Surprisingly, we find that preference optimization methods consistently improve TruthfulQA performance compared to the SFT checkpoint, and the improvement could be as high as over 10% in some cases. Similarly, we hypothesize that the preference dataset contains instances that emphasize truthfulness, which helps the model better understand the context and generate more truthful responses.
Math performance drops.
GSM8K is the benchmark that shows the most volatility across methods. Notably, except for ORPO, almost all approaches lead to consistent drops in one or more settings. We hypothesize that ORPO retains performance largely due to its supervised fine-tuning loss for regulation. adds a reference-model calibrated supervised fine-tuning loss to the preference optimization objective, and find that it effectively solves the issue and maintains performance on math tasks as well.
Overall, identifying a pattern in downstream performance is challenging. Comprehensive analysis is difficult due to using different pretrained models, preference optimization datasets, and objectives. Recent works indicate that gradient-based approaches could be effective in finding relevant data for downstream tasks , and could possibly extended to understand the effect of preference optimization. We believe a thorough study on how preference optimization affects downstream performance would be valuable and call for a rigorous and more comprehensive analysis in future work.
Appendix D Standard Deviation of AlpacaEval 2 and Arena-Hard
We present the standard deviation of AlpacaEval 2 and the 95% confidence interval of Arena-Hard in Table 10. All these metrics are reasonable and do not exhibit any significant outliers or instability.
Appendix E Generation Length Analysis
Removing length normalization from the SimPO objective results in an approach similar to Contrastive Preference Optimization (CPO) , which interpolates reward maximization with a supervised fine-tuning loss and has demonstrated strong performance in machine translation. However, without the supervised fine-tuning loss, the reward maximization objective without length normalization is suboptimal in preference optimization.
We analyze the generation length of models trained with or without length normalization on AlpacaEval 2 and Arena-Hard. As shown in Figure 6, length normalization significantly decrease the generation length by up to 25% compared to when it is not used in most cases. However, even though the generation length is shorter, the models with length normalization consistently achieve much higher win rates on both benchmarks. This suggests that length normalization can effectively control the verbosity of the generated responses, and meanwhile improve the generation quality.
Length is not a reliable indicator of generation quality.
We further analyze the generation length of models trained with different methods on AlpacaEval 2 and Arena-Hard, as shown in Table 10. Generally, we find that no single method consistently generates longer or shorter responses across all settings. Additionally, even though some methods may generate longer responses, they do not necessarily achieve better win rates on the benchmarks. This indicates that the length of the generated responses is not a reliable indicator of generation quality.
SimPO demonstrates minimal exploitation of response length.
We observe that SimPO has a shorter generation length compared to DPO in the Llama-3-Instruct case but exhibits a higher generation length in other settings, with up to 26% longer responses on AlpacaEval 2. Conversely, SimPO only increases length by only around 5% on Arena-Hard compared to DPO. It is fair to say that the generation length heavily depends on the evaluation benchmark. A stronger indicator is that SimPO consistently achieves a higher length-controlled win rate on AlpacaEval 2 compared to the raw win rate, demonstrating minimal exploitation of response length.
Appendix F Gradient Analysis
We examine the gradients of SimPO and DPO to understand their different impact on the training process.
represent the gradient weight in SimPO and DPO, respectively. It can be seen that the differences are twofold: (1) comparing the gradient weights and , SimPO’s gradient weight does not involve the reference model and has a straightforward interpretation: the weights will be higher for samples where the policy model incorrectly assigns higher likelihood to than ; (2) comparing the gradient updates, SimPO’s gradients on and are length-normalized, while DPO’s are not. This corresponds to the empirical findings that DPO may exploit length bias: longer sequences with more tokens will receive larger gradient updates in DPO, dominating the training process.
Appendix G Qualitative Analysis
We present the win rate heatmap of Mistral-Base and Mistral-Instruct on AlpacaEval 2 and Arena-Hard in Figure 7 and Figure 8, respectively. Based on this analysis, we present qualitative examples of responses generated by a SimPO model, a DPO model and the baseline model GPT-4-Preview-1106 on AlpacaEval 2.
In Figure 9 and Figure 10, we present an example where Mistral-Base-SimPO generates a better-structured answer compared to Mistral-Base-DPO. Given the question, "How can you determine if a person is genuinely interested in a conversation or simply being polite?", the DPO model generates a response with a long list of bullet points, making it difficult to understand the relationships between different points. In contrast, the SimPO model produces a well-structured answer with high-level categorization of different behaviors, followed by detailed suggestions for each category. This makes the answer more readable and easier to understand.
Comparing Instruct models with Base models when trained with SimPO.
In Figure 11, we present an example where Llama-3-Instruct generates a more detailed and well-formatted answer compared to the baseline model, and as well as the Llama-3-Base-SimPO model. Given the question: What language does Argentina people speak? Llama-3-Base-SimPO only gives a very brief answer. GPT-4-Preview-1106 gives a more detailed answer in explaining how the Argentina Spanish differs from standard Spanish. However, the answer is not well formatted and a bit hard to parse. Llama-3-Instruct-SimPO gives a detailed and well-formatted answer, which is easier to read and understand, and offers sufficient details.
Appendix H Llama-3-Instruct v0.2 (Jul 7, 2024)
In this section, we update the Llama-3-Instruct setting, primarily by utilizing a stronger reward model to annotate our generated preference dataset.
In our previous version, we use PairRM as our reward model to rank generated candidate responses. The results, presented in Table 12, show that switching the reward model from PairRM to ArmoRM for ranking the data markedly improves model performance. This underscores the importance of a high-quality preference optimization dataset for enhancing performance. Notably, SimPO has achieved a 53.7 LC win rate on AlpacaEval 2 and 36.5 on Arena-Hard, surpassing the previous version by 9.0 and 2.7 points, respectively.
We use the following hyperparameters for SimPO under the Llama-3 Instruct v0.2 setting: and . The other hyperparameters (e.g., learning rate, batch size, max sequence lengths) are kept the same as the original Llama-3-8B-Instruct setting.
Strong SFT model and high-quality policy data diminish algorithm differences.
With a strong SFT model like Llama-3-8B-Instruct, and as the preference optimization data quality improves, the differences between algorithms become less pronounced. For instance, DPO achieved a similar win rate as SimPO in terms of raw win rate, and DPO, IPO, and R-DPO all exhibited comparable raw win rates on Arena-Hard. However, SimPO maintains an advantage by producing shorter sequences, resulting in a significantly better LC win rate on AlpacaEval 2.
Stronger downstream task performance.
The v0.2 version also shows improved performance in downstream tasks across various objectives. However, DPO, IPO, R-DPO, and SimPO continue to experience a decline in reasoning-intensive domains such as GSM8K. In contrast, objectives that include an SFT component maintain their performance in mathematical tasks.
Incorporating SFT regularization in SimPO.
Several reference-free algorithms, including RRHF , SLiC-HF , CPO , and ORPO , employ SFT regularization in their objectives. SFT regularization can be an effective method to prevent reward hacking, ensuring that the solution maintains low loss without resulting in degraded generations. We also experiment with the integration of an SFT loss in SimPO, yielding the following objective:
As shown in Table 14, the addition of the SFT regularization leads to a decrease in performance on AlpacaEval 2. However, we note that SFT regularization provides substantial benefits to certain tasks such as GSM8K, as shown in Table 12. These contrasting results suggest that the impact of SFT in preference optimization may vary depending on the training setup and the nature of the task. Further comprehensive studies on this topic are left for future research.
Appendix I Applying Length Normalization and Target Reward Margin to DPO (Jul 7, 2024)
Since the release of the paper, we have had inquiries from researchers about whether the key design elements of SimPO—length normalization and target reward margin—could benefit DPO. By doing so, we will derive the following two objectives:
An intuitive understanding of how length normalization could benefit DPO is that, despite DPO’s reward design being implicitly normalized by the reference model, the policy model might still exploit length bias from the data, resulting in a disproportionately high probability for longer sequences. Applying length normalization could help mitigate this effect.
We train models with the objectives mentioned above and compare their performance to that of DPO and SimPO, as shown in Table 15.
The results indicate that, unlike SimPO, length normalization and target reward margin do not consistently benefit DPO. Specifically, length normalization significantly improves DPO performance only in the Mistral-Base setting, where the preference optimization dataset shows a strong length bias. However, it does not provide a benefit in the Mistral-Instruct setting, where the lengths of winning and losing responses are comparable. This is likely because DPO already includes an implicit instance-wise target reward margin via the reference model, as shown in the derivation below.
Appendix J Applying SimPO to Gemma 2 Models (Sept 16, 2024)
After releasing the Llama-3-SimPO checkpoints, we received extensive feedback about performance degradation on benchmarks measuring specific capabilities, such as MMLU and GSM8K. To investigate this issue, we continued training the Llama-3-8B-Instruct model with different learning rates, as reported in Table 16. We find that using a higher learning rate results in a stronger model in chat-oriented benchmarks, at the cost of catastrophic forgetting on GSM8K and MMLU. We evaluate the zero-shot performance of the models on GSM8K and MMLU using the ZeroEval repository which adopts a unified setup. With a smaller learning rate, the model’s performance on chat benchmarks is slightly worse, but its performance on GSM8K and MMLU is better retained. This demonstrates a trade-off between chat-oriented benchmarks and other benchmarks when continuing training from a strong instruction-tuned model.
Applying SimPO to Gemma 2 models presents a different trend.
We evaluate SimPO using Google’s recently released Gemma-2-9B-it model , which represents a strong open-source model. For training data, we generate up to 5 responses per prompt from the UltraFeedback dataset and use the ArmoRM model to annotate preferences between responses. We compare our SimPO against a DPO-trained variant, both fine-tuned from the Gemma-2-9B-it base model. As shown in Table 17, SimPO demonstrates superior performance on chat benchmarks like AlpacaEval 2 and Arena-Hard while maintaining the model’s original zero-shot capabilities on tasks like GSM8K and MMLU. Notably, we find that varying the learning rate during fine-tuning has minimal impact on the model’s performance. These results suggest an underlying property difference between the Llama-3 checkpoints and the Gemma 2 checkpoints, and might be worth further investigation.
Gemma-2-9B-it-SimPO significantly improved the ranking of the Gemma-2-9B-it model on Chatbot Arena.
During the development stage, we relied solely on automated metrics to evaluate the model’s performance. To determine if these metrics aligned with real user preferences, we submitted our best-performing model, Gemma-2-9B-it-SimPO, to the Chatbot Arena leaderboard hosted by LMSYS . We find that our model improved the original Gemma-2-9B-it ranking from 36th to 25th, making the SimPO variant the top-ranked 10B model on the Chatbot Arena leaderboard based on real user votes as of September 16th, 2024.