Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging
Joel Jang, Seungone Kim, Bill Yuchen Lin, Yizhong Wang, Jack Hessel, Luke Zettlemoyer, Hannaneh Hajishirzi, Yejin Choi, Prithviraj Ammanabrolu
Introduction
Reinforcement Learning from Human Feedback (RLHF) (Nakano et al., 2021a; Ouyang et al., 2022a; Bai et al., 2022a; Dubois et al., 2023; Bai et al., 2022b) typically optimizes a policy model that receives training signals from a single reward model that aims to capture the general preferences of a population. In this work, we instead propose Reinforcement Learning from Personalized Human Feedback (RLHF), a new, multi-objective formulation of the human preference alignment problem, where Large Language Models (LLMs) are trained to be efficiently aligned with a range of different, potentially personalized combinations of human preferences.
We model RLHF as a Multi-Objective Reinforcement Learning (MORL) problem, which allows training the policy model with multiple, conflicting objectives since it aims to vary the importance of each objective during inference. In existing RLHF formulations, pairwise human feedback is collected by asking human annotators to choose which model response is generally better and is used to train a general reward model. This makes implicit assumptions that may not hold for everyone. For example, recent work has shown that LLMs aligned with RLHF prefer verbose output generations (Zheng et al., 2023; Dubois et al., 2023; Wang et al., 2023; Singhal et al., 2023). We aim to support a wider range of multifaceted preferences that are explicitly declared as desirable by the user—giving the user control over the facets of output text they want to see as well as the personal data they wish to reveal to the model. We collect personalized human feedback corresponding to multiple such dimensions, noting that they may also be conflicting in nature.
We first implement a strong MORL baseline called Prompted-MORL where there are multiple reward signals for each of the objectives (preferences) given via prompts during RL training. Next, we propose Personalized Soups, a method that circumvents simultaneously optimizing multiple preferences by first optimizing multiple policy models each with distinct preferences with Proximal Policy Optimization (PPO) and merging the parameters of the policy models whose preferences we want to composite together on the fly during inference. This modular approach significantly reduces the computational complexity from exponential to linear in relation to the total number of unique preferences. Furthermore, since Personalized Soups does not have to be trained in a multitask fashion, it does not require re-training the underlying policy every time a novel preference (objective) is added.
We empirically show that by transforming the problem of aligning LLMs to human preferences into a MORL problem, we are able to have personalized alignment that provides a deeper level of adaptation to individual users that supervised fine-tuning, RLHF, and prompting cannot attain. We also emphasize the modularity of Personalized Soups by performing experiments in a scenario where the user additionally writes novel preferences that they want to integrate with existing preferences. We show that in this scenario, Personalized Soups still performs competitively to Prompted-MORL while being exponentially more efficient through parameter merging.
Related Work
Incorporating human preference feedback into a reward model, and subsequently optimizing a language model to output text that reward model scores highly with an RL algorithm, has been shown to result in language models that generate outputs humans generally prefer (Ouyang et al., 2022b). This process has been applied to summarization (Ziegler et al., 2019; Stiennon et al., 2020; Wu et al., 2021a), answering questions with long-form answers using text retrieved from the web (Nakano et al., 2021b; Menick et al., 2022), generating engaging responses in a dialogue settings (Thoppilan et al., 2022; Cohen et al., 2022) and following human instructions (Kojima et al., 2021; Suhr & Artzi, 2022; Kim et al., 2023).
However, the standard RLHF setup commonly addressed in prior work assumes a reward model that accounts only for average annotator preference, i.e., the fact that different users may desire different outputs, even for the same prompt, is ignored Casper et al. (2023). Individual preferences can vary not only on aesthetic axes, but also on semantics. For example, Santurkar et al. (2023) use public opinion polling to show that “default” LLM preferences vary in their degree of expressed-opinion alignment with different average opinions among demographic groups.Feng et al. (2023) suggests that “default” LLM expressed opinions stem directly from the pretraining data. Kirk et al. (2023) defines a taxonomy and policy framework for the alignment of LLMs with personalized feedback. While Wu et al. (2023) performs fine-grained RLHF which is very similar in spirit and allows personalization, our work develops MORL algorithms for scenarios where there are conflicting preferences, not only orthogonal objectives.
In this work, we propose formulating LLM personalization as a MORL problem, which was typically studied in decision-making tasks (Hayes et al., 2022) that aims to tackle the problem of simply optimizing by a single, scalar, additive reward function (Sutton & Barto, 2018), which possesses many limitations such as (1) suboptimal solutions due to lack of representation (Hayes et al., 2022), (2) lack of explainability of distinct objectives, and (3) ensuring fair outcomes for multiple participants (Vamplew et al., 2018; Siddique et al., 2020).
Previous work has aimed to alleviate these problems through novel MORL methods (Van Moffaert et al., 2013; Van Moffaert & Nowé, 2014; Yang et al., 2019; Xu et al., 2020). Other work aims to solve complex problems such as water management, military purchasing, wind farm control, etc. (Hayes et al., 2022) by converting the single-objective RL problem into a MORL problem. In this work, we convert the problem of aligning LLMs to human preferences into a MORL problem to (1) provide a more optimal solution for each individual, (2) allow users to dynamically choose the distinct objectives they want to optimize, and (3) ensure fairness by allowing preferences that may be in the long-tail to be integrated.
Personalization in Natural Language Processing (NLP) has mainly been focused on creating personalized dialogue agents (Zhang et al., 2018; Mazaré et al., 2018; Zheng et al., 2019; Wu et al., 2021b; Xu et al., 2022), where the task is to create chitchat agents that are engaging with distinct personas based on user profile (e.g. gender, age, residence, etc.) or past user history data (e.g. Reddit posts, etc.). Another line of work (Salemi et al., 2023) leverages personalized information to boost performance on specific tasks such as review generation (Li & Tuzhilin, 2019), recipe generation (Majumder et al., 2019), and headline generation (Ao et al., 2021). This line of work requires model providers to make better models utilizing the personal information of the user. In our work, we propose a framework that allows users to choose which preference the language model should prefer, essentially giving control to the user.
Recent work has shown that performing weighted linear interpolation of model parameters leads to the composition of each model ability (Li et al., 2022; Wortsman et al., 2022b; a; Don-Yehiya et al., 2022; Huang et al., 2023). This line of work has led to many interesting applications of model merging such as composing the abilities of expert models that perform different tasks (Ilharco et al., 2022; Jang et al., 2023) and introducing language-specific modules for growing the total capacity of multilingual LMs (Pfeiffer et al., 2022).
Most recently, Rame et al. (2023) proposed to merge policy models that were trained to perform specific tasks such as question answering and summarization using proxy reward models. While they mostly deal with reward models trained on the same data, our proposed MORL methods are an extension of this work that actually deals with diverse reward models trained on multifaceted human feedback to show compositional abilities through parameter merging rather than just ensembling.
Reinforcement Learning from Personalized Human Feedback
The current RLHF can be denoted as optimizing policy :
where is the reward model trained on general human feedback. As pointed out in Silver et al. (2021), the following optimization may implicitly be occurring under the hood:
where represents rewards from objectives that human annotators may generally consider positive objectives (e.g., informative, helpful, kind, etc.) and is the total number of unique ‘dimensions’ of these positive objectives. This formulation does not allow modeling conflicting objectives, which may occur in real-world scenarios. For example, some people may prefer concise and unpretentious responses in contrast to informative, polite responses.
In this section, we formalize RLHF where we allow modeling conflicting preferences during alignment. We explain how we collect conflicting feedback in Section 3.1. In Section 3.2, we explain how we convert the current RLHF formulation into a MORL problem. Lastly, we explain the details of our evaluation in Section 3.3.
We utilize Tulu-7B LM (Wang et al., 2023), a model that uses LLaMA-7B (Touvron et al., 2023) as a base model and is instruction tuned on a mixture of open-source instruction-tuning datasets, as the base model for our experiments. We utilize 10k prompt instances from GPT4-Alpaca (Peng et al., 2023), one of the datasets used to train Tulu-7B, as our instruction dataset to generate rollouts and collect pairwise feedback data. We also use the same during Proximal Policy Optimization (PPO) training (Schulman et al., 2017) of Tulu-7B.
Following previous work, we simulate human annotators with GPT-4 for collecting large-scale pairwise feedback data (Bai et al., 2022b; Dubois et al., 2023)—but note that our evaluations are validated with (smaller-scale) human preference data collected from crowdworkers. While Dubois et al. (2023) mostly simulates GPT-4 and other LLMs to choose which is generally a better response between two candidate responses, we provide GPT-4 with a single preference (full list shown in Table 1) to decide which is a better response. We also provide the same preference criteria via additional prompts during the rollout generation of the two candidate responses; we use Tulu-30B for the rollout generation while the actual policy model we train is Tulu-7B for our main experimental setup, making our experimental setting an off-policy training set-up.
While we have feedback on which of the two model responses is more aligned with a single preference via GPT-4 annotation, utilizing only two positive pairs during reward model training was empirically shown to be less robust during the PPO training. Instead, we train our reward model on multiple comparisons (Song et al., 2023; Kim et al., 2023) by including a neutral response and a negative response as shown in Figure 2. Specifically, the reward model is provided with four different comparisons for a single prompt during training: positive 1 positive 2 (decided by GPT-4), positive neutral, positive negative, and neutral negative. The positive response when compared with the neutral and the negative response is chosen randomly. This allows the reward model to be exposed to different granularity of the specific preference and give scores accordingly during PPO training. We explore (1) training a single reward model in a multitask fashion that leverages the preference prompts during inference to give distinct rewards according to each preference and (2) training multiple reward models, each tailored to the distinct preference.
2 Multi-Objective Reinforcement Learning (MORL)
where is the policy model that aims to maximize multiple objectives from the rewards and is the importance placed on each objective. If are constants during PPO training, this problem essentially becomes a single objective problem, maximizing towards a single, general objective. In our setup, we have conflicting preferences which require dynamically varying in a binary manner with respect to the conflicting preference during training and inference.
First, we introduce a strong baseline that varies during MORL through prompts. While are given as inputs directly to the policy model in traditional RL settings using PPO with MORL, there is no straightforward way of integrating different as an input to LLMs. Instead, we utilize the preference prompts as binary signals for .
We append the unique preference combination (shown in Table 1) with a training prompt from { + P1 + P2 + P3} P1, P2, P3 each represents preference prompts from each preference dimension in Table 1. For one example, one unique combination might be P1A + P2B + P3A (ABA) where the combined objective for the response needs to be elementary level, informative, and friendly. before feeding it to our initial policy model and getting the output response. Then, we gather reward signals for each of the preferences by feeding { + P1/P2/P3 + output} into a single reward model (doing three forward passes) Empirically, utilizing a single reward model instead of multiple reward models led to better performance. We hypothesize this is due to the problem of normalizing signals from different reward models (Hayes et al., 2022), which is known to be a nontrivial problem in MORL. to get the reward signal specific to the individual preference and averaging the three reward values to get a single scalar reward used for PPO training. We multitask train the policy model across the eight different unique combinations of preferences, which essentially results in varying .
While Prompted-MORL can be a clear baseline for converting the alignment problem into a MORL problem, we propose another approach that does not have to see all existing preference combinations during training thus allowing increasing the total number of distinct preferences at scale, which is required for true personalization.
We decompose the MORL problem into multiple single-objective problems:
where we optimize each policy individually. Then during inference, we pick and choose the policy models whose objective we want to maximize together and perform a weighted sum of the parameters on the fly:
This means that even though the exact preference combinations haven’t been seen during training, we are still able to composite them on the fly during inference. While the total computational complexity increases exponentially when we are required to observe all possible combinations, optimizing individual objectives separately only increases the complexity in linear space.
This also means that multitask training is not necessary and allows efficient integration of novel preferences. Since personalization also entails that there can be an infinite number of new preference dimensions, we assert that Personalized Soups makes tackling RLHF feasible.
3 Multifaceted Evaluation
For evaluation, we manually filter out 50 instances from the Koala evaluation (Geng et al., 2023) that require open-ended generations. We also modified some of the prompts so that the evaluation prompts do not contain any elements requiring individual preferences (e.g., removing the phrase asking for a elementary-level response from the original prompt since we want to test the LLM to generate a expert-level response). The full list of evaluation prompts used for our experiments is shown in Appendix C. In our evaluation setup, we simulate users to have a unique combination of preferences, each from the three preference dimensions (Expertise, Informativeness, Style) in Table 1, which equates to 8 unique preference combinations (examples shown in Figure 1). We get the average win rate across the simulated 8 preference combinations for our final evaluation. We use a variant of the AlpacaFarm evaluation framework for simulated (GPT4) evaluation and hire 24 crowdworkers for human evaluation. Details of human evaluation are provided in Appendix A.
When given an evaluation prompt , we first generate responses from each model to get the outputs where is model A and is model B. The common approach is get {Win, Tie, Lose} by asking the human which model response is generally preferred.
In our evaluation setup, we first assign scores to each of the possible feedback: Win = 1, Tie = 0, Lose = -1. Next, we iterate through the different preference dimensions and get an aggregated score value: . Finally, we have Win if , Tie if , and Lose if . To get the final win rate between vs. , we iterate through the entire evaluation set (50 prompts) the unique preference combinations (8 combinations) and get the total as the final win rate, while disregarding the total number of ties.
Experiments
In this subsection, we provide details of the single-objective baseline methods we implement. The summary of the key component differences in comparison with our proposed methods is provided in Table 2.
Vanilla Baseline (VB) As the most simple baseline, we simply utilize the base Tulu-7B model to generate responses without providing it any notion of preferences. During the evaluation, we use the same response to evaluate on the 8 different preference combinations.
Reinforcement Learning from Human Feedback (RLHF) We perform RLHF in the traditional manner where GPT-4 labels which response is generally better, train a reward model using the pairwise feedback data, and use the reward model to adapt the policy model with PPO training. The same 10k instances from are used for RLHF.
Preference Prompting (PP) Next, we observe how far the instruction-tuned base LM can integrate multiple preference combinations by simply prompting for the preferences without any additional training.
Multi-task Training (MT) For a competitive baseline, we utilize the positive candidate selected by GPT-4 as the output for imitation learning, which is essentially performing rejection sampling (Nakano et al., 2021a) that uses GPT-4 as the reward model for selecting golden responses from the distribution of responses. We append the individual preference prompt with instances from and multitask train the Tulu-7B model across all six individual preferences. This method also allows distilling the outputs of Tulu-30B for training Tulu-7B.
2 Experimental Details
For both the reward model and policy training, we limit ourselves to going through only once (1 epoch). In the initial exploration stage, the end performance for the policy model did not improve even if we trained the reward model for longer epochs. For policy model training, we utilize our evaluation dataset (50 prompts) to get the average reward and chose the policy model checkpoint that showed the highest average reward on the evaluation set for our final evaluation. We utilize LoRA (Hu et al., 2022) for both the reward model and policy model training. The detailed hyperparameters for the reward model and policy model training are provided in our github repository https://github.com/joeljang/RLPHF.
3 Main Results
Table 3 and 4 show the results of doing all possible pairwise comparisons across the methods using GPT-4 and humans as judges, respectively. Note that the win rate of each battle is calculated using the aggregated win rate explained in Section 3.3. Each individual preference combination results are shown in Appendix D. We also show the average criteria-wise win rate instead of the aggregated win rate across all of the methods in Appendix B.
The first thing to note is that there is a limitation to the extent prompting (PP) can integrate multiple preferences. This means that specific training for integrating the multiple preferences is necessary to composite them together. Next, we can see that supervised fine-tuning (RS) underperforms MORL-based methods, which is consistent with prior work that also showed the advantage of RL-based approaches when aligning LLMs with human feedback compared to its supervised-finetuning counterpart. Finally, while P-Morl and P-Soups both outperform other methods on average, there exists a discrepancy between the simulated and human evaluation; P-Soups has the highest average win rate in GPT-4 evaluation while P-Morl has the highest in human evaluation. Nonetheless, P-Soups is able to show superior performance in comparison to baseline methods and competitive performance to P-Morl.
In previous parameter merging literature, multi-task fine-tuning (RS) used to be considered the upper bound for compositional abilities through parameter merging (Don-Yehiya et al., 2022). However, in our scenario, we can see that parameter merging (P-Soups) is able to outperform multitask fine-tuning, showing promising results for parameter merging not only as a distributed multitask finetuning method but a method that can result in superior performance than multitask training.
One might still wonder about the general ‘helpfulness’ capabilities of models that are trained to be tailored to multiple preferences. In Figure 3, we first show the average pairwise win rate from Table 3 in green. Next, we instead ONLY perform pairwise comparisons with the unseen objective ‘helpfulness’ (ask GPT-4 which model response they generally prefer better) and report the average pairwise win rate in red.
RLHF performs the best in this scenario, which shows that there is no free lunch; while the objective of RLHF was to provide model responses that are generally preferred (highly correlated with ‘helpfulness’), the other methods were prompted/trained to be optimized towards the personalized preference aspects, possibly deviating away from general helpfulness. While RS, P-Morl, and P-Soups are able to retain similar performance in terms of helpfulness compared to the initial instruction-tuned model (VB), we observe that prompting (PP) significantly underperforms compared to other methods which also highlights the limitation of simply prompting base/instruction-tuned models for personalized preferences and shows the need for specialized training methods for personalization.
4 Scaling to New Preferences
While we explore 6 distinct preferences in this work, we are still limited in doing ‘declarative’ personalization; that is, the individual preferences have been pre-defined to measure the performance of different methodologies. However, in the real world, individuals may not be bound by pre-defined preferences. Furthermore, people’s preferences might change over time, which requires continual learning of new preferences. This means that we may be required to train infinite numbers of preferences to be truly personalized to individuals’ preferences. Considering this aspect, the scalability of methods becomes a critical factor in implementing RLPHF in real-world scenarios.
In order to compare the scalability of P-Morl and P-Soups, we add two new preferences (in addition to the ones in Table 1 to the Style dimensions: (P3C) “Generate/Choose a response (that answers) in a sassy manner.” and (P3D) “Generate/Choose a response (that answers) in a sarcastic manner.”, which results in a total of 16 (2 2 4) unique preference combinations. We re-train P-Morl on the 16 new preference combinations and only train two new policy models for integrating P-Soups. The simulated win-rate between P-Morl and P-soups on each of the original preference combinations (53.04% win rate of P-Soup over P-Morl in Table 3 decomposed into each preference combinations) and the 16 new preference combinations are shown in Figure 4.
As shown in the figure, P-Soups shows competitive performance compared to P-Morl while being much more efficient considering that it (1) did not have to observe all 16 possible preference combinations and (2) did not have to re-train on the previous preferences, but just train two new policies each for the new preference in a modular manner and merge their parameters on-the-fly during inference. Considering that P-Morl is bounded by while P-Soups is bounded by where is the total number of preferences (assuming there are two unique preferences for each dimension), we assert that P-Soups allows tackling RLPHF to be feasible.
Conclusion
Previous work has shown that adapting LLMs with RLHF helps them generate outputs that are preferred by humans over the supervised fine-tuned counterpart. However, recent work has also pointed out that simply training LLMs to abide by the preference of the general may result in ignoring individual preferences and values. In this work, we provide the first steps to tackle this issue by proposing Reinforcement Learning from Personalized Human Feedback as a multi-objective problem so that LLMs can be aligned to follow conflicting preferences. We propose a promising method called P-Soups that is able to composite models trained on single objectives on the fly during inference. We also highlight the scalability of P-Soups by showing that it scales linearly, instead of exponentially like the MORL baseline, with regards to the number of new preferences, which is required to provide true personalization to individual users.
Thanks to Minyoung Hwang, Sungdong Kim, Tim Dettmers, Yoonjoo Lee, and Margaret Li for helpful feedback and discussion.
References
Appendix A Details of Evaluation Setup
We use a modified version of the GPT4 annotation prompt used by Dubois et al. (2023). We modify the criteria to perform the pairwise evaluation from general to a single preference dimension. We also provide 3 demonstrations: one scenario where there is a tie because both responses do not contain any notion of the preference (e.g. both responses do not show any signs of friendliness), one scenario where there is a clear winner, and one scenario where they are both good, but one is better than the other.
We recruited 24 crowd workers for our human evaluation. Figure 5 shows the interface used for human evaluation. We consider both the ‘Tie’ and ‘Both are bad’ options to be Ties.
Appendix B Criteria-wise Evaluation
The criteria-wise win rate (%) across all the methods are shown in Figure 6. The criteria-wise win rate was calculated by getting the average win-rate of the preference combinations that contained the specific preference dimension. For example, when calculating the criteria-wise win rate of ‘Elementary’, we got the average win-rate of the preference combinations that contained the ‘Elementary’ preference, which includes AAA, AAB, ABA, and ABB.
Appendix C The full list of evaluation prompts
The full list of evaluation prompts used in our experiments are provided in Table 5.
Appendix D Detailed results for Human Evaluation and GPT-4 Evaluation
We provide detailed results (win / loss / tie for each of the preference combinations for our main experimental results. Table 6 shows the GPT4 evaluation and Table 7 shows the human evaluation results.
Appendix E Examples of P-Soups text generations
Table 8 shows empirical examples of the text generated from each preference combination of the 16 preference combination experiments for the same prompt.