Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Haoxiang Wang, Yong Lin, Wei Xiong, Rui Yang, Shizhe Diao, Shuang Qiu, Han Zhao, Tong Zhang

Introduction

Large language models (LLMs) (OpenAI, 2023; Anthropic, 2023) have demonstrated remarkable capabilities across various domains and tasks, such as mathematical reasoning (Wei et al., 2022) and medical question answering (Singhal et al., 2023a; Wang et al., 2023a; Thirunavukarasu et al., 2023). However, for an assistant to be truly useful, it must align with human preferences, such as being helpful, honest, harmless, and managing verbosity.

Reinforcement Learning from Human Feedback (RLHF) (Christiano et al., 2017; Ziegler et al., 2019; Ouyang et al., 2022; Bai et al., 2022b; Lee et al., 2023), is the leading approach to adapt LLMs towards these complex, often implicitly-defined goals. Typically, the most popular RLHF framework (Christiano et al., 2017; Ziegler et al., 2019; Ouyang et al., 2022) first constructs a scalar reward model to represent the difficult-to-specify goal of being preferred by human and then use this reward model to provide signals for the subsequent reward optimization stage. Its success spans various practical applications, including recommendation systems (Pereira et al., 2019), image generation (Hao et al., 2022; Wu et al., 2023a; Dong et al., 2023a), robotics (Brown et al., 2019), and most notably, aligning LLMs with human values and preferences, such as ChatGPT (OpenAI, 2023), Claude (Anthropic, 2023), Llama 2 (Touvron et al., 2023) and Gemini (Team et al., 2023).

While recent advancements in RLHF are noteworthy, a fundamental challenge persists due to problem misspecification. This means that a single reward function may not sufficiently capture complex human values. For example, a generative model aligned by RLHF for helpfulness tends to produce verbose responses as shown in Figure 1 (Left) Singhal et al. (2023b), even though many users prefer answers that are both helpful and concise. Assuming ae-objective reward implies a total order over preferences, which is hard to satisfy when the preference is aggregated across a diverse set of human groups (May, 1954; Tversky, 1969), because humans typically have a set of intricate or even contradictory targets (Biyik and Sadigh, 2018). In real-world applications, the scalar-reward RLHF tends to align the LLMs toward an “average-user” preference, which cannot capture the complicated nature of human preferences and can be unfair for the under-represented groups (Feffer et al., 2023). For example, consider User-1, 2, 3, and responses AA, BB, CC in Fig. 2 (Left). User-1 and 3 prefer response BB over CC (B≺CB\prec C), while User-2 prefers CC over BB (C≺BC\prec B). This could occur as response CC is more verbose than BB, while User-2 prefers concise answers. When these diverse preferences are aggregated across human groups, the typical reward models with scalar rewards tend to learn the “average-user” preference (which is B≺CB\prec C in this case), overlooking the individual preference of User-2, as shown in Figure 2 (Middle). This is also known as the “Condorcet paradox” in the theory of social choice Gehrlein (2002). In general, human opinions and expertise can vary significantly (Coello, 2000; Bobu et al., 2023; Bansal et al., 2023). Meanwhile, the importance of these targets may also change over time, depending on the users and their expectations.

To address the limitations of the existing scalar reward model, previous works suggest the use of multi-objective rewards that characterize human preferences from different aspects (e.g., helpfulness, verbosity, harmlessness) (Pan et al., 2023; Rame et al., 2023). One common way is to take the human feedback as a multi-dimensional reward vector and each dimension models one objective (Rame et al., 2023; Dong et al., 2023b). Then, one may apply a linear combination to transform the multi-objective rewards into a scalar for LLM alignment (Bakker et al., 2022; Wu et al., 2023b). However, this approach still cannot handle the user-dependent needs from a diverse user population and can be unfair for minority groups. One may further adopt a user-dependent linear combination to multi-objective rewards for aligning a model for each user preference (Rame et al., 2023; Jang et al., 2023). However, this approach is quite inference-unfriendly because we have to switch between different models in response to the different user preferences. Finally, in social choice theory, a game-based formulation was studied under the name maximal lotteries (Sternberg, 1965; Fishburn, 1984), as well as the subsequent works in RLHF (Wang et al., 2023b; Swamy et al., 2024; Ye et al., 2024), to handle the diversity of user preferences. We remark that their framework is fundamentally different from the multi-objective rewards and cannot offer a user-dependent preference control in the inference stage, either. Refer to Section 2.3 for a more detailed discussion with existing methods.

In recognition of the aforementioned limitations, we propose a novel and practical alignment approach, Directional Preference Alignment (DPA), to enhance the adaptability and controllability of a single LLM. Our aligned LLM enjoys the flexibility to be controlled with different preferences embedded numerically into the system prompt. The ability to control preferences can significantly enhance the model’s personalization ability during inference. For example, as the model is aligned with DPA with helpfulness and verbosity in consideration, a user could simply control the model’s generation by specifying a directional preference v=<v1,v2>v=\left<v_{1},v_{2}\right> that ∥v∥2=1\|v\|_{2}=1, and the model will generate responses that maximize reward=v1×helpfulness+v2×verbosity\texttt{reward}=v_{1}\times\texttt{helpfulness}+v_{2}\times\texttt{verbosity} where helpfulness and verbosity are rewards scored from different perspectives as shown in Figure 1 (Right). Figure 2 (Right) further shows that the preferences of User-1, User-2, and User-3 can be accurately represented by specifying the preference vector in the 2-dimensional space. This is a scenario where DPA can alleviate the problem of misspecification in RLHF.

Our approach features two crucial aspects: 1). Multi-Objective Rewards, which involve learning with multiple different preference targets simultaneously, and 2). Directional Preference Alignment, which encodes user preferences as unit vectors for preference-aware LLM alignment. Specifically, we summarize our contributions as follows.

We identify the limitations of existing popular RLHF frameworks: 1) the limited capacity for capturing the real-world complicated human preference; 2) lacking in adaptability for user-dependent preference;

We propose Directional Preference Alignment (DPA): a novel alignment approach that allows a single LLM to accommodate users with varying preferences.

We consider both helpfulness and verbosity rewards, and align Mistral-7B (Jiang et al., 2023) with our DPA: empirical evaluations show that DPA offers effective arithmetic control over the trade-off between helpfulness and verbosity, while maintaining competitive performance with DPO (Rafailov et al., 2023).

Directional Preference Alignment

In a typical RLHF pipeline (Ouyang et al., 2022; Bai et al., 2022a; Touvron et al., 2023), we first construct a reward model based on a labeled preference dataset (e.g., preference A≺B≺CA\prec B\prec C annotated by a labeler) and then use the reward model to provide supervision for the subsequent reward optimization stage. In this section, we first present the problem setup, where we additionally consider multi-objective rewards and user preferences in the framework. Then, we present our algorithm, the Directional Preference Alignment, to handle the problem of preference-aware alignment.

1 Multi-Objective Reward Model

We consider kk-objective reward for a response yy given prompt xx as

2 Directional Preference Alignment

Our work aims to learn a collection of policies that can traverse the Pareto front as efficiently as possible. Moreover, we intend to relate the learned policies to the user’s preferences concerning various objectives and control the learning process according to such preferences. To make multi-objective optimization tractable and controllable, a common approach is linear scalarization (Caruana, 1997; Ghane-Kanafi and Khorram, 2015; Hu et al., 2023), which takes a linear combination of multiple objectives. Through exploring all different linear combinations, the solutions to these problems can sufficiently cover a significant area of the Pareto front, which justifies the application of the linear scalarization approach.

To incorporate user preference into the language model, we condition the text generation on vv in addition to xx, such that the response is generated according to y∼πθ(⋅∣x,v)y\sim\pi_{\theta}(\cdot|x,v). For a specific vv, the preference-conditional reward objective is

Reward Optimization via Rejection Sampling.

We now proceed to discuss the algorithmic designs for optimizing the RL objective in Eq. (4). While PPO is the most predominant approach for a fixed reward function (OpenAI, 2023; Anthropic, 2023), it is known that PPO is unstable and sample-inefficient in aligning LLMs (Choshen et al., 2019) and imposes a heavy burden on GPU memory resources (Ouyang et al., 2022; Yuan et al., 2023). Hence, PPO requires extensive efforts to be tuned to its best performance. In light of the above limitations, we resort to an alternative approach, Rejection Sampling Fine-tuning (RSF) (Dong et al., 2023a; Yuan et al., 2023; Gulcehre et al., 2023), a RLHF algorithm used in the Llama 2 project (Touvron et al., 2023), with appealing simplicity, stability, and comparable reward gains. In essence, the original RSF learns from the best-of-nn policy created by the reward function. Initially, we generate nn responses using a base LLM and then rank them using the reward model to select the responses with the highest reward. We further finetune our LLM based on these selected samples, and this process can be repeated multiple times.

In our scenario, to address the multi-objective nature and user-dependent preferences, we iteratively alternate among the following steps for t=1,…,Tt=1,\dots,T iterations:

Preparation. Initialize an empty dataset Dt=∅\mathcal{D}_{t}=\emptyset. Prepare policy model πθt−1\pi_{\theta_{t-1}} obtained from last iteration.

The whole procedure of our methods is summarized in Figure 3.

3 Discussion with Existing Methods

Comparison with SteerLM Dong et al. (2023b). Recall that we have multi-objective reward r=<r1,r2,...,rk>r=\left<r_{1},r_{2},...,r_{k}\right> of each response yy to the prompt xx. Dong et al. (2023b) first fine-tunes the generative model to maximize the likelihood of yy by taking both xx and rr as the input prompts:

When presented with a new input xˉ\bar{x}, SteerLM aims to produce a response that aligns with the newly assigned multi-dimensional rˉ\bar{r}. Particularly, a user could specify rˉ\bar{r} as “(helpfulness=10,verbosity=1)"(\texttt{helpfulness}=10,\texttt{verbosity}=1)", namely high helpfulness but low verbosity, for a new prompt xˉ=\mbox‘‘Pleasesummarize‘RomeoandJuliet′"\bar{x}=\mbox{``Please summarize `Romeo and Juliet'"}. SteerLM could then generate answers according to rˉ\bar{r}. However, SteerLM will encounter a significant challenge when a user-specified rˉ\bar{r} falls outside the feasible region of rewards for the given xˉ\bar{x}, i.e., rˉ∉{r:(xˉ,y,r)∈Dr}\bar{r}\notin\{r:(\bar{x},y,r)\in\mathcal{D}_{r}\}. In this case, if a user sets a rˉ\bar{r} that is not achievable given xˉ\bar{x}, SteerLM may generate uncontrolled responses due to the infeasibility of rˉ\bar{r} under xˉ\bar{x}. For example, “(helpfulness=10,verbosity=1)(\texttt{helpfulness}=10,\texttt{verbosity}=1)” could be infeasible for xˉ\bar{x} according to the set S\mathcal{S} since it will be difficult or impossible to generate a helpful summarization of ‘Romeo and Juliet’ in very few words.

Comparison with Soup Methods Rame et al. (2023); Jang et al. (2023). Soup methods trains a policy θi\theta_{i} for each reward objective. Let ri(x,y)r_{i}(x,y) denote the ii-th objective, we have:

Empirical Results

We conduct experiments on Mistral-7B (Jiang et al., 2023), focusing on two reward objectives: helpfulness and verbosity. Our proposed DPA achieves arithmetic control of LLM generations for different helpfulness-verbosity preferences while demonstrating an excellent balance between the two objectives.

Recently, the verbosity bias in LLMs and humans, meaning that LLMs and humans sometimes prefer more verbose answers even though they are of similar qualities, has attracted considerable attention (Saito et al., 2023; Singhal et al., 2023b). It has been exploited or even “hacked” by the RLHF-aligned models. For instance, Kabir et al. (2023) demonstrated that 77%77\% of ChatGPT answers are verbose, while Yuan et al. (2024) found that the average output length increases to 2.5 times as the DPO iterates. Preliminary experiments have been conducted in response to this bias, such as those by Chen et al. (2024), which explicitly consider verbosity as a response feature. Benchmark creators like AlpacaEval (Li et al., 2023) and MT-Bench (Zheng et al., 2023) have observed verbosity bias in their LLM judges (typically GPT-4), and AlpacaEval-2.0 has adjusted to account for output lengthtatsu-lab.github.io/alpaca_eval/.

1 Implementation

We use two datasets for experiments: HelpSteer and UltraFeedback. Both datasets are used for reward model trainingWe include HelpSteer since it has verbosity annotations., while only UltraFeedback is used for finetuning.

HelpSteer Wang et al. (2023d) comprises 10K prompts and 37K annotated responses with five attributes: helpfulness, correctness, coherence, complexity, and verbosity. A 43B closed-source LLM generated responses, and human labelers annotated each response on a scale of 0-4 for the five attributes.

UltraFeedback Cui et al. (2023) includes 64K prompts, each of them are associated with 4 responses of five attributes: honesty, truthfulness, instruction-following, helpfulness and overall-score. GPT-4 was employed to label these responses. We use the same training-validation prompt split hf.co/datasets/HuggingFaceH4/ultrafeedback_binarized as Zephyr (Tunstall et al., 2023).

Reward Modeling.

We train a multi-objective reward model on the union of HelpSteer and UltraFeedback, initializing with Mistral-7B. Specifically, we follow SteerLM-v2 practicesThe authors of SteerLM (Dong et al., 2023b) improved the original training recipe in a follow-up work (Wang et al., 2023c), which we denote as SteerLM-v2. (Wang et al., 2023c), attaching a linear regression head layer on the last hidden state of Mistral-7B. We include both regression and traditional language modeling losses in the reward model training, as we find the latter improves accuracy without additional observed costs. The reward model has 10 output dimensions: the first half corresponds to HelpSteer’s five attributes, while the other half accounts for UltraFeedback’s attributes. Rewards in each dimension are rescaled to the range of 0-100 in the data preprocessing stage.

Alignment Setup.

For a fair comparison with DPO (Rafailov et al., 2023), we conduct a head-to-head comparison with Zephyr-β\beta (Tunstall et al., 2023), a DPO-trained Mistral-7B model that was state-of-the-art (7B) at its release. Zephyr-β\beta uses supervised fine-tuning (SFT) on UltraChat-200K (Ding et al., 2023) followed by DPO on UltraFeedback (Cui et al., 2023). Since RLHF typically begins with SFT models, we initialize with the SFT checkpoint of Zephyr-β\beta and apply DPA on UltraFeedback. Following practices of Cui et al. (2023); Tunstall et al. (2023), we average instruction-following, truthfulness, honesty, and helpfulness ratings of UltraFeedback for the overall helpfulness objective. We use HelpSteer’s verbosity attribute for the verbosity objective. Our multi-objective reward model annotates helpfulness and verbosity for all UltraFeedback data and self-generated responses.

Rewards and Directional Preferences.

Dataset Splitting.

Iterative RLHF methods typically sample responses for unseen prompts in each new iteration to prevent the model from simply memorizing and repeating the responses (Dong et al., 2023a; Xiong et al., 2023; Yuan et al., 2024). In view of this, we split UltraFeedback dataset into two disjoint subsets, D1\mathcal{D}_{1} and D2\mathcal{D}_{2}, containing an equal number of unique prompts. In each iteration tt, we initialize the policy model πθt\pi_{\theta_{t}} from an SFT checkpoint rather than πθt−1\pi_{\theta_{t-1}}, and we use a different subset from the last iteration. The use of alternative subsets ensures that the policy model πθt\pi_{\theta_{t}} for response sampling in iteration t+1t+1 has not encountered the prompts before.

Rejection Sampling.

We conduct rejection sampling following our iterative algorithm detailed in Sec. 2.2. Notice that to launch training in t=1t=1, we need πθt=0\pi_{\theta_{t=0}} for sampling responses for a diverse set of helpfulness-verbosity preferences. However, Zephyr-β\beta-SFT is not designed for preference-conditional generation, making it not a good choice for πθt=0\pi_{\theta_{t=0}}. To resolve this, we train a SteerLM model on D2\mathcal{D}_{2} (a half of UltraFeedback) that can generate responses conditioned on both user prompt xx (sampled from D1\mathcal{D}_{1}) and reward objectives r1,r2r_{1},r_{2}. We use this model for rejection sampling in iteration t=1t=1 to obtain πθ1\pi_{\theta_{1}} (for each prompt, we generate 80 responses for diverse reward combinations (r1,r2)(r_{1},r_{2})). In all the following iterations, for each prompt, we sample 5 directional preferences <v1,v2>\left<v_{1},v_{2}\right>, and use πθt−1\pi_{\theta_{t-1}} to generate 16 responses per preference, then keep the highest-reward response and reject the rest 15.

Fine-tuning.

For the response data obtained through rejection sampling, we prepend the user’s directional preference to the system prompt, as illustrated in Fig. 1, to make the model aware of the user preference. The fine-tuning process then follows the same approach as SFT, optimizing the next-token prediction loss across the text corpus. It is also worth noting that RLHF often leads to performance degradation or knowledge forgetting, a phenomenon referred to as alignment tax in the literature (Askell et al., 2021; Lin et al., 2023). To mitigate this issue, we adopt the memory replay techniques suggested in Instruct-GPT (Ouyang et al., 2022) and Llama 2 (Touvron et al., 2023) that can effectively reduce alignment tax (Lin et al., 2023). Specifically, we incorporate original responses from UltraFeedback, which constitute about 15% of our finetuning data for each iteration. Our algorithm is applied for iterations t=1,…,4t=1,\dots,4.

Software, Hardware and Hyperparameters

We use PyTorch (Paszke et al., 2019) with HuggingFace’s TRL framework (von Werra et al., 2020) for all fine-tuning experiments across t=0,…,Tt=0,\dots,T. All experiments are conducted on 8x A6000 GPUs. The training cost of each DPA iteration is about 60 GPU hours. The AdamW optimizer (Loshchilov and Hutter, 2019) is employed with a learning rate of 10−510^{-5} and a cosine learning rate schedule (20 warmup steps). We use a context window of 4096 tokens with sample-packing (packing short responses within the context window). The training takes 2 epochs with a global batch size of 64. We use vLLM (Kwon et al., 2023) for inference. In the rejection sampling process, we conduct inference with temperature 1.01.0. In evaluation (Sec. 3.2), we use temperature 0.70.7.

2 Evaluation

For validation, we used 2000 prompts from UltraFeedback and considered 10 uniformly sampled directional preferences ranging from v=<1,0>v=\left<1,0\right> to v=<2/2,2/2>v=\left<\sqrt{2}/2,\sqrt{2}/2\right>. For each prompt-preference combination, our DPA-aligned models generated two responses. We then calculated the average helpfulness and verbosity rewards for all 2000 responses per preference using our reward model. For SteerLMWe trained a SteerLM model (initialized with the SFT checkpoint of Zephyr-β\beta) on UltraFeedback, following practices of Wang et al. (2023c)., five verbosity reward values were sampled, and the highest corresponding helpfulness reward from UltraFeedback was identified for each value. These verbosity-helpfulness pairs were then used to condition SteerLM’s generation, with the average rewards computed across prompts. In the case of Zephyr-β\beta’s DPO and SFT models, we generated responses using their original prompt templates and averaged the rewards across the validation set. The results, illustrated in Fig. 4, show that as t≥1t\geq 1, our DPA model Pareto-dominates SFT, DPO, SteerLM, and DPA at iteration tt Pareto-dominates the models of previous iterations. This demonstrates DPA’s effective arithmetic control for different user preferences, and with increasing finetuning iterations tt, the empirical front of DPA (i.e., each curve in Fig. 4) expands, indicating that our finetuning approach successfully maximizes rewards for all user preferences of consideration. Notably, our DPA’s empirical front significantly surpasses that of SteerLM and DPO, even though all models were trained on the same UltraFeedback dataset and originated from the same SFT model.

AlpacaEval-2.0 Evaluation

AlpacaEval-2.0 (Li et al., 2023) is an LLM-based automatic evaluation benchmark that employs GPT-4-turbo as the LLM judge. It includes 805 prompts, and model responses to these prompts are compared with reference answers provided by GPT-4-turbo. Subsequently, the win-rate against the reference answers is calculated as a metric for the models’ instruction-following capabilities. We evaluated SteerLM and our DPA (at t=4t=4) conditioned with various user preferences and report the win rate and average response length in Fig. 5, along with DPO and SFT results for reference. Fig. 5 demonstrates that our DPA model outperforms SteerLM and achieves competitive performance against DPO while providing arithmetic control for diverse user preferences. The discrepancy between the validation reward evaluation results and the AlpacaEval-2.0 outcomes may arise because our reward model has different behaviors and preferences compared to GPT-4-turbo. While DPA can closely fit the reward model, this does not necessarily guarantee generalization to GPT-4-turbo evaluations.

Related Works

The landscape of natural language processing has been profoundly transformed in recent years through the development of large language models (LLMs), showcasing human-level proficiency across a range of tasks including text classification, generation, and complex reasoning. This progress stems from extensive pre-training on vast datasets, enabling these models to address diverse challenges. Despite their achievements, a distinction arises between closed-source models (e.g., GPT-3 (Brown et al., 2020), Bard (Google, 2023), Claude (Anthropic, 2023), and PaLM (Chowdhery et al., 2023)), often surpassing their open-source counterparts (e.g., megatron-turing-530b (Smith et al., 2022), and Bloom (Workshop et al., 2022)) in performance (Liang et al., 2022), which poses challenges for open-source research. However, initiatives like Meta’s LLaMA (Touvron et al., 2023) and subsequent works such as Alpaca (Taori et al., 2023), Vicuna (Chiang et al., 2023), and LMFlow (Diao et al., 2023), demonstrate significant open-source contributions that continue to push the boundaries of what’s possible with LLMs. These advancements enabled by the fine-tuning techniques, aim to improve LLMs’ ability and adapt to a wide range of domains and tasks. Nonetheless, as these generative foundation models advance, they still face problems like implicit biases, underscoring the need for ongoing alignment and ethical considerations in their development and application. In this paper, we focus on how to align LLMs with human preferences, including the principles of being helpful, honest, and harmless as outlined by (Askell et al., 2021). This procedure is often achieved by Reinforcement Learning with Human Feedback (RLHF) Ouyang et al. (2022).

RLHF Algorithmic Designs.

Policy Optimization (PPO) (Schulman et al., 2017) is the most predominant approach, with its tremendous success in Chat-GPT (OpenAI, 2023) and Claude (Anthropic, 2023). However, PPO is significantly less efficient and stable compared to supervised finetuning (Choshen et al., 2019), and is also sensitive to the parameter and code-level implementation (Engstrom et al., 2020). Therefore, tuning the PPO to its best performance is very challenging in practice and the results of Chat-GPT (OpenAI, 2023) have not been widely reproduced so far. In view of this, efforts have been made to develop supervised-learning-based methods as an alternative approach to the PPO, and we review them as follows. Rejection sampling finetuning (RSF) is proposed in (Dong et al., 2023a; Yuan et al., 2023; Gulcehre et al., 2023) with different variants, but essentially, they learn from the positive samples selected by a learned reward model. RSF was applied to the RLHF of LLaMA2 project (Touvron et al., 2023) and we adopt the iterative implementation as suggested in Dong et al. (2023a); Touvron et al. (2023); Gulcehre et al. (2023). There is also another line of work designing algorithms from the KL-constraint reward optimization (Rafailov et al., 2023; Zhao et al., 2023; Azar et al., 2023; Xiong et al., 2023), which additionally requires the resulting model to be close to the initial model. Among them, the Direct Preference Optimization (DPO) (Rafailov et al., 2023) has attracted considerable attention due to its simplicity and stability, and effectiveness. We remark that it is also possible to incorporate these algorithmic ideas into our DPA framework and we leave the algorithmic design beyond RSF to future work.

Fine-grained Preference Representation and Algorithmic design.

The scalar-reward-model has been criticized mainly due to its limited capacity (Wu et al., 2023b; Casper et al., 2023; Munos et al., 2023) (see the discussion of preference intransitivity in Section 1 for an illustrative example). A line of works has considered multi-objective rewards to capture the different aspects of human preferences (Zhou et al., 2023; Jang et al., 2023; Touvron et al., 2023; Wu et al., 2023b; Köpf et al., 2023; Rame et al., 2023). However, the multi-objective rewards are then combined in a fixed way (e.g., Wu et al., 2023b; Touvron et al., 2023), mainly to represent a preference averaged over different human groups, failing to capture the user-dependent preference. By introducing the user preference as a unit vector (direction) into the directional preference alignment framework, we achieve a fine-grained and user-dependent representation for the complicated human preference. Notably, in social choice theory (Sternberg, 1965; Fishburn, 1984), as well as some very recent studies in RLHF (Wang et al., 2023b; Swamy et al., 2024; Ye et al., 2024), the RLHF is formulated as a game between two LLMs to partially handle the diversity of preferences in the population-level. The learning objective is accordingly adjusted to be solving the Nash equilibrium of the game. In comparison, our techniques are fundamentally different from theirs and may offer computational advantages since game-based formulation is far more complicated.

Limitations

A primary constraint of our DPA framework is its reliance on a robust multi-objective reward model. The efficacy of DPA is intrinsically linked to the precision and discriminative capability of this reward model. Should the reward model not adequately capture the subtleties of specific preferences or exhibit bias in its reward distribution, the DPA might inadvertently exacerbate these shortcomings throughout the fine-tuning process. Furthermore, if the reward model fails to recognize harmful content, it could lead the aligned model to produce such content during inference.

Conclusion

In this paper, we introduce Directional Preference Alignment (DPA) to incorporate multidimensional user preferences. DPA addresses the limitation of conventional scalar reward models by alleviating conflicting user preferences through a high-dimensional preference vector in a multidimensional space. We demonstrate that DPA efficiently explores the Pareto front in the multidimensional reward space, revealing a more effective trade-off between helpfulness and verbosity on Mistral-7B compared to existing strong baselines such as DPO.

References