PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward Model

Baijiong Lin, Weisen Jiang, Yuancheng Xu, Hao Chen, Ying-Cong Chen

Introduction

The alignment of large language models (LLMs) is crucial to ensure that their outputs reflect human values (Wang et al. 2023; Casper et al. 2024). In practice, human preferences and values are often multifaceted and may conflict. For example, users may expect LLM responses to be simultaneously helpful, harmless, and humorous. These competing objectives pose a challenge for single-objective alignment methods to meet such complex demands. To address this, multi-objective alignment enables LLMs to dynamically adjust trade-offs among different preference dimensions based on specific user needs (represented as a preference vector).

Current multi-objective alignment methods (Zhou et al. 2024; Rame et al. 2023; Jang et al. 2023; Wang et al. 2024a; Guo et al. 2024; Yang et al. 2024b; Zhong et al. 2024) require extensive computations for LLM training that many researchers and practitioners cannot access when the LLM is large (e.g., with 65B parameters). Different from existing methods, we focus on multi-objective test-time alignment that trains a small reward model rather than the original LLM, reducing computation cost largely and making multi-objective alignment accessible with limited computing resources.

GenARM (Xu et al. 2025), the recent state-of-the-art test-time alignment method, introduces an Autoregressive Reward Model (ARM) to predict token-level rewards for guiding the generation of frozen LLMs during inference, resulting in effective and efficient test-time alignment. However, when adapted to multi-objective settings, GenARM requires training an ARM for each objective. During inference, each ARM computes its respective next-token reward, and the LLM’s generation is guided by a weighted sum of these rewards using the given preference vector as weights. Hence, GenARM faces two main limitations in multi-objective test-time alignment: 1. the need for multiple ARMs generation increases the inference cost, and 2. the separate training of ARMs leads to potential misalignment between the guided generation and the specified preference vector.

To address these limitations, we propose preference-aware ARM (PARM) to achieve effective and efficient multi-objective test-time alignment. Unlike GenARM, which independently trains ARMs for each preference dimension without awareness of other dimensions, PARM is a single unified model jointly trained across all preference dimensions to explicitly optimize trade-offs between different preferences. Moreover, PARM is conditioned on preference vectors, allowing it to dynamically adjust the output reward according to the user-specific preference vector during inference, thereby guiding the frozen LLM to generate responses that align with the given preference vector while maintaining computational efficiency.

To condition the PARM (which may contain billions of parameters) on a low-dimensional preference vector, we propose preference-aware bilinear low-rank adaptation (PBLoRA). PBLoRA employs a bilinear form BWA{\mathbf{B}}{\mathbf{W}}{\mathbf{A}}, where B{\mathbf{B}} and A{\mathbf{A}} are low-rank matrices, similar to those used in LoRA (Hu et al. 2022). W{\mathbf{W}} is an r×rr\times r weight matrix (rr is the rank) that is conditioned on the preference vector. This conditioning allows the preference vector to directly control the generation of PARM through W{\mathbf{W}}. Moreover, we theoretically show that the bilinear form used in PBLoRA is more expressive than the original LoRA, enabling PARM to better capture the complex relationships between different preference dimensions. By training with PBLoRA, PARM is steerable to output reward according to the user-specific preference vector, thereby guiding the frozen LLM in generating responses aligned with the given preference vector during inference.

We evaluate PARM on the safety alignment (Ji et al. 2023; Ji et al. 2024) and helpful assistant (Bai et al. 2022) tasks. Experimental results demonstrate that PARM has higher alignment quality and is more inference-efficient than previous methods in multi-objective test-time alignment. Moreover, we highlight the weak-to-strong guidance ability of PARM, where a smaller PARM can guide a larger frozen LLM (e.g., 7B guides 65B) without the need for training the larger LLM, making multi-objective alignment accessible with limited computing resources.

The overall comparison between PARM and existing multi-objective alignment methods is shown in Table 1. As shown, PARM only needs to train a single small reward model rather than the original LLM or multiple reward models. This significantly reduces computation cost, facilitating multi-objective alignment under computational constraints.

The contributions of this paper are summarized as follows: 1. We propose PARM, a single unified ARM jointly trained across all preference dimensions, to achieve effective and efficient multi-objective test-time alignment; 2. We propose PBLoRA to adapt the ARM to condition on the preference vector, enabling better management of trade-offs between preferences; 3. Experiments show that PARM significantly reduces inference cost while improving alignment performance compared with existing methods. Moreover, PARM enables weak-to-strong guidance, aligning larger LLMs with a smaller PARM, eliminating the need for expensive training of the larger models.

Related Work

Test-Time Alignment. Let πbase{\bm{\pi}_{\text{base}}} denote a base model. As directly fine-tuning πbase{\bm{\pi}_{\text{base}}} on preference data to align with the human values (Ouyang et al. 2022; Rafailov et al. 2023; Meng et al. 2024; Park et al. 2024) requires extensive computation on LLM training, test-time alignment methods keep πbase{\bm{\pi}_{\text{base}}} frozen and use reward models to guide the generation during inference. Existing test-time alignment methods are inspired by the closed-form solution of RLHF in (Rafailov et al. 2023) as follows,

where π{\bm{\pi}} is the aligned model, Z(x)Z({\mathbf{x}}) is the partition function, r(x,y)r({\mathbf{x}},{\mathbf{y}}) is a reward model, and β\beta is a hyperparameter. According to Equation (1), when the base model πbase{\bm{\pi}_{\text{base}}} is frozen, the generation is guided by the reward model r(x,y)r({\mathbf{x}},{\mathbf{y}}).

When generating the next token from an incomplete response based on Equation (1), it is necessary to predict rewards for the next token. However, such token-level rewards cannot be directly obtained from response-level reward models. Some test-time methods, like (Khanov et al. 2024; Li et al. 2024), attempt to compute rewards using incomplete responses, which often leads to inaccuracies. Alternatively, methods such as (Huang et al. 2024; Chakraborty et al. 2024) generate complete responses to compute rewards for each token, which significantly increases inference costs.

Recently, GenARM (Xu et al. 2025) proposes the Autoregressive Reward Model (ARM), which explicitly predicts token-level rewards to enable efficient and effective test-time alignment. However, when extending to multi-objective scenarios that need to handle trade-offs among multiple preference dimensions, GenARM has two significant limitations (introduced in Section 4.1). Hence, in this paper, we focus on improving the ARM to enable efficient and effective multi-objective test-time alignment. Similarly, PAD (Chen et al. 2025a) also employs a token-level reward model to guide the decoding process. However, it focuses on aligning with personalized preferences rather than managing trade-offs across different preference dimensions.

Multi-Objective Alignment. In practice, human preferences are multi-dimensional and we often need to align LLMs to balance multiple, sometimes conflicting, preference dimensions such as helpfulness, harmlessness, and humor (Yang et al. 2024b). Some multi-objective alignment methods like (Wu et al. 2023; Zhou et al. 2024) train separate LLMs for each given preference vector by linearly combining multiple reward models. To reduce the training cost, some methods separately train specialized LLMs for each preference dimension and integrate either through parameter fusion (Rame et al. 2023; Jang et al. 2023) or logit combination (Shi et al. 2024) during inference. However, this strategy still requires maintaining multiple models, resulting in significant storage and computational burdens. To further improve efficiency, some methods focus on adapting a single LLM to accommodate varying preferences. This is achieved through encoding preference vectors into input prompts (Wang et al. 2024a; Guo et al. 2024; Yang et al. 2024b) or directly modifying model parameters (Wang et al. 2024b; Zhong et al. 2024).

However, existing multi-objective alignment methods typically require direct fine-tuning of base LLMs, incurring substantial computational costs. In this paper, we focus on multi-objective test-time alignment, where base LLMs remain frozen, eliminating the need for expensive fine-tuning.

We also review multi-objective optimization and controllable text generation in Appendix A.

Preliminary on ARM

In this section, we review the recent test-time alignment method, GenARM (Xu et al. 2025), which uses autoregressive reward models (ARM) to guide the generation of frozen LLMs during inference.

ARM. ARM is a token-level reward model, whose reward r(x,y)r({\mathbf{x}},{\mathbf{y}}) is computed as the sum of log probabilities of tokens generated up to the tt-th token, as follows,

where πθ(⋅∣x,y<t){\bm{\pi_{\theta}}}(\cdot|{\mathbf{x}},{\mathbf{y}}_{<t}) is a learnable distribution function (parameterized by θ{\bm{\theta}}) that predicts the next-token reward. Most practical language model architectures are autoregressive, e.g., the LLaMA family of models (Touvron et al. 2023), thus, can be employed for πθ{\bm{\pi_{\theta}}}.

Training of ARM. Let D={(x,y1,y2,z)}{\bm{\mathcal{D}}}=\{({\mathbf{x}},{\mathbf{y}}^{1},{\mathbf{y}}^{2},z)\} denote a preference dataset, where y1{\mathbf{y}}^{1} and y2{\mathbf{y}}^{2} represent the different responses generated by πbase{\bm{\pi}_{\text{base}}} in response to the prompt x{\mathbf{x}}, and z∈{0,1}z\in\{0,1\} is the preference label (z=1z=1 if y1{\mathbf{y}}^{1} is a “better” response than y2{\mathbf{y}}^{2} otherwise 00). The ARM is trained on D{\bm{\mathcal{D}}} using a negative log-likelihood loss function as follows,

where σ(⋅)\sigma(\cdot) is the logistic function and βr\beta_{r} is a hyperparameter.

Guided Generation via ARM. GenARM (Xu et al. 2025) achieves test-time alignment by integrating the trained ARM into Equation (1) as

The probability of the next token yty_{t} is conditioned on the partially generated response y<t{\mathbf{y}}_{<t} and prompt x{\mathbf{x}} as follows,

which resembles decoding from multiple language models, enabling us to leverage prior methods such as (Dekoninck et al. 2024).

PARM: Preference-Aware ARM

In this section, we introduce the preference-aware ARM (PARM) for multi-objective test-time alignment.

Multi-objective test-time alignment aims to use reward models to guide the base model to generate a response that aligns with multi-dimensional user preferences during inference while keeping the base model frozen.

Let D={(x,y1,y2,z1,⋯ ,zk)}{\bm{\mathcal{D}}}=\{({\mathbf{x}},{\mathbf{y}}^{1},{\mathbf{y}}^{2},z_{1},\cdots,z_{k})\} denote a kk-dimensional preference dataset, where zi=1z_{i}=1 if y1{\mathbf{y}}^{1} is a “better” response than y2{\mathbf{y}}^{2} in the ii-th preference dimension otherwise 00. We denote Di={(x,y1,y2,zi)}{\bm{\mathcal{D}}}_{i}=\{({\mathbf{x}},{\mathbf{y}}^{1},{\mathbf{y}}^{2},z_{i})\} as the preference dataset for the ii-th dimension. In multi-objective alignment, users expect the LLM’s outputs to align with their multi-dimensional needs, which can be represented as a preference vector, α=(α1,⋯ ,αk)∈Δk−1{\bm{\alpha}}=(\alpha_{1},\cdots,\alpha_{k})\in{\Delta_{k-1}}, where αi\alpha_{i} denotes the weight for the ii-th preference dimension and Δk−1={α∣∑i=1kαi=1,αi≥0,i=1,⋯ ,k}{\Delta_{k-1}}=\{{\bm{\alpha}}|\sum_{i=1}^{k}\alpha_{i}=1,\alpha_{i}\geq 0,i=1,\cdots,k\} is a (k−1)(k-1)-dimensional simplex.

The recent state-of-the-art method GenARM (Xu et al. 2025) first trains an ARM πθi{\bm{\pi}_{{\bm{\theta}}_{i}}} on dataset Di{\bm{\mathcal{D}}}_{i} for each preference dimension ii. During inference, given a preference vector α{\bm{\alpha}}, the separated trained ARMs are then combined to guide the generation procedure as,

GenARM faces two major limitations in multi-objective test-time alignment: 1. The kk ARMs are trained independently on different preference dimensions without awareness of each other, causing potential conflicts when combining their rewards directly during inference (i.e., Equation 6), resulting in a mismatch between model outputs and the desired preferences. 2. GenARM needs kk ARMs to predict reward simultaneously in the generation process, causing a huge computational overhead during inference.

To address these limitations, we aim to jointly train a single ARM across all preferences by optimizing the following multi-objective optimization problem,

To learn the whole Pareto optimal solutions in a single run, we propose to train a unified ARM conditioning on the preference vector, i.e., θ(α){\bm{\theta}}({\bm{\alpha}}), called the Preference-aware ARM (PARM). This conditioning enables a single ARM to approximate the whole Pareto set and effectively manage trade-offs across different preference dimensions. Therefore, given a preference vector α{\bm{\alpha}} at inference, we can obtain the corresponding Pareto-optimal ARM without retraining and use this ARM to guide the frozen base LLM to generate responses aligned with the preference, thereby addressing the misalignment and inefficiency issues of GenARM.

In the following, we introduce how to condition the ARM on preference vectors in Section 4.2, how to train PARM in Section 4.3, and how to guide generation via PARM for multi-objective test-time alignment in Section 4.4.

2 Preference-Aware Bilinear Low-Rank Adaptation

Similar to GenARM (Xu et al. 2025), we adopt the autoregressive model for the reward model πθ(α)(⋅∣x,y<t){\bm{\pi}_{{\bm{\theta}}({\bm{\alpha}})}}(\cdot|{\mathbf{x}},{\mathbf{y}}_{<t}). The primary challenge is how to condition the massive model parameters θ{\bm{\theta}} (which may contain billions of parameters) on the kk-dimensional preference vector α{\bm{\alpha}}.

In this paper, we propose preference-aware bilinear low-rank adaptation (PBLoRA) for PARM, enabling efficient and effective conditioning on preference vectors while maintaining computational scalability. Low-rank adaptation (LoRA) (Hu et al. 2022) is a widely used parameter-efficient technique for fine-tuning LLMs. However, the simple product of two low-rank matrices fails to account for user preferences. To address this issue, we propose Preference-Aware Bilinear Low-Rank Adaptation (PBLoRA) to condition on preference vectors as follows.

Assume both B{\mathbf{B}} and A{\mathbf{A}} have rank rr. The outer product of {bi ⁣: ⁣i ⁣= ⁣1,⋯ ,r}\{{\mathbf{b}}_{i}\!:\!i\!=\!1,\cdots,r\} with {ai ⁣: ⁣i ⁣= ⁣1,⋯ ,r}\{{\mathbf{a}}_{i}\!:\!i\!=\!1,\cdots,r\} results in r2r^{2} linearly independent matrices {biaj⊤ ⁣ ⁣ ⁣: ⁣i ⁣= ⁣1,⋯ ,r,j ⁣= ⁣1,⋯ ,r}\{\mathbf{b}_{i}\mathbf{a}_{j}^{\top}\!\!\!:\!i\!=\!1,\cdots,r,j\!=\!1,\cdots,r\} in the space of m ⁣× ⁣nm\!\times\!n matrices.

The proof is provided in Appendix C. The bilinear form BWA=∑i=1r∑j=1rwijbiaj⊤{\mathbf{B}}{\mathbf{W}}{\mathbf{A}}=\sum_{i=1}^{r}\sum_{j=1}^{r}w_{ij}{\mathbf{b}}_{i}{\mathbf{a}}_{j}^{\top} is in the subspace spanned by {biaj⊤ ⁣: ⁣i ⁣= ⁣1,⋯ ,r,j ⁣= ⁣1,⋯ ,r}\{\mathbf{b}_{i}\mathbf{a}_{j}^{\top}\!:\!i\!=\!1,\cdots,r,j\!=\!1,\cdots,r\}, while the original LoRA formulation BA=∑i=1rbiai⊤{\mathbf{B}}{\mathbf{A}}=\sum_{i=1}^{r}{\mathbf{b}}_{i}{\mathbf{a}}_{i}^{\top} is in the subspace spanned by {biai⊤ ⁣: ⁣i ⁣= ⁣1,⋯ ,r}\{\mathbf{b}_{i}\mathbf{a}_{i}^{\top}\!:\!i\!=\!1,\cdots,r\}. According to Theorem 4.1, the former subspace has r2r^{2} dimensionality and is rr times higher than the latter (only rr). Hence, the formulation BWA{\mathbf{B}}{\mathbf{W}}{\mathbf{A}} is more expressive than BA{\mathbf{B}}{\mathbf{A}} in the space of m×nm\times n matrices.

The term BW(α)A{\mathbf{B}}{\mathbf{W}}({\bm{\alpha}}){\mathbf{A}} in Equation 8 is preference-aware since W{\mathbf{W}} is conditioned on the preference vector α{\bm{\alpha}}. As multiple objectives may share common knowledge (Zhang & Yang 2022; Chen et al. 2025b), we split BW(α)A{\mathbf{B}}{\mathbf{W}}({\bm{\alpha}}){\mathbf{A}} into two terms: a preference-agnostic term to learn shared features and a preference-aware one to learn objective-specific features, as follows,

The preference-agnostic term B1W1A1{\mathbf{B}}_{1}{\mathbf{W}}_{1}{\mathbf{A}}_{1} is shared among different α{\bm{\alpha}} and thus can explicitly learn shared features across different preference dimensions. Meanwhile, the preference-aware term B2W2(α)A2{\mathbf{B}}_{2}{\mathbf{W}}_{2}({\bm{\alpha}}){\mathbf{A}}_{2} captures the specific adjustments required for each unique preference vector, enabling fine-grained alignment with individual objectives.

Moreover, PBLoRA is parameter-efficient. The total parameter size of (m+n)×(r1+r2)+r12+kr22≈(m+n)×(r1+r2)(m+n)\times(r_{1}+r_{2})+r_{1}^{2}+kr_{2}^{2}\approx(m+n)\times(r_{1}+r_{2}), since k,r1,r2≪{m,n}k,r_{1},r_{2}\ll\{m,n\}. This suggests that PBLoRA can handle kk preference dimensions using almost the same number of parameters compared to LoRA with rank r1+r2r_{1}+r_{2}. Compared with GenARM, which requires training kk ARMs (implemented by LoRA with rank r1+r2r_{1}+r_{2}), PBLoRA is roughly k×k\times more parameter-efficient.

3 Training of PARM

At the training stage of PARM, we keep θ0{\bm{\theta}}_{0} frozen and only learn the parameters Θ={A1,A2,B1,B2,W1,ϕ}{\bm{\Theta}}=\{{\mathbf{A}}_{1},{\mathbf{A}}_{2},{\mathbf{B}}_{1},{\mathbf{B}}_{2},{\mathbf{W}}_{1},{\bm{\phi}}\} for PBLoRA. The training objective of PARM is formulated as follows,

4 Guided Generation via PARM

The trained PARM is used to guide the autoregressive generation of the frozen base LLM πbase{\bm{\pi}_{\text{base}}} under any user-specific preference vector α{\bm{\alpha}}.

Given an α{\bm{\alpha}}, we compute the reward of PARM as

According to Equation 17, we compute the next-token conditional probability as follows,

Unlike GenARM (Xu et al. 2025), which relies on kk ARMs to compute rewards (i.e., Equation 6), PARM operates with a single unified reward model, contributing to k×k\times faster inference.

Experiments

In this section, we evaluate PARM through experiments on safety alignment and helpful assistant tasks, demonstrating its effectiveness and efficiency in multi-objective test-time alignment. Our implementation is based on the open-source trl library (von Werra et al. 2020).

Experimental Setups. Safety alignment aims to balance the helpfulness and harmlessness in language models when responding to red-teaming prompts. We use the PKU-SafeRLHF-10K dataset (Ji et al. 2023; Ji et al. 2024), which provides harmlessness and helpfulness annotations for each question-answering (QA) pair. Following (Zhou et al. 2024), we randomly split the dataset into three parts: 88K samples for training, 0.50.5K for validation, and the remaining 1.51.5K for testing. Following (Zhou et al. 2024), we use two open-source pretrained reward models from (Ji et al. 2023) as oracles to score the harmlessness and helpfulness for each response, respectively. Following GenARM (Xu et al. 2025), we employ the Alpaca-7B model (Taori et al. 2023) as the base model πbase{\bm{\pi}_{\text{base}}}. Both the ARMs in GenARM (Xu et al. 2025) and our PARM are fine-tuned from the Alpaca-7B model. The sources of dataset and models are provided in Appendix F.

Baselines. We compare the proposed PARM with the following baselines: 1. Rewarded soups (RS) (Rame et al. 2023) that fine-tunes kk base models and weights them as a single model at the parameter space using the given preference vector α{\bm{\alpha}} for inference; 2. MOD (Shi et al. 2024) that fine-tunes kk base models and combines their logits using the given preference vector α{\bm{\alpha}} at inference; 3. GenARM (Xu et al. 2025) that trains kk ARMs while keeping the base model frozen and uses the trained ARMs to guide the generation of the frozen base model.

Implementation Details. The proposed PARM is fine-tuned from the Alpaca-7B model using PBLoRA for 22 epochs with βr=0.01\beta_{r}=0.01, a learning rate of 5×10−45\times 10^{-4}, and a total batch size of 3232. Our implementation is based on the peft library (Mangrulkar et al. 2022), where PBLoRA is applied to the query, key, and value weight matrices in the attention layers. Both r1r_{1} and r2r_{2} in PBLoRA are set to 44.

For the baseline GenARM (Xu et al. 2025), two separate ARMs are trained for helpfulness and harmlessness, respectively, using the same training settings. Specifically, we fine-tune the Alpaca-7B model with LoRA (Hu et al. 2022) for 11 epoch, employing βr=0.01\beta_{r}=0.01, a learning rate of 5×10−45\times 10^{-4}, and a total batch size of 3232. LoRA with a rank of 88 is applied to the same layers as PBLoRA.

For the baselines RS (Rame et al. 2023) and MOD (Shi et al. 2024), two separate DPO models (Rafailov et al. 2023) are fine-tuned from the Alpaca-7B model using LoRA (Hu et al. 2022) for helpfulness and harmlessness, respectively, using the same training settings as GenARM.

During generation, we set β=1\beta=1 and use a maximum generation length of 10241024 tokens for all methods.

Evaluation. We evaluate all methods on the test dataset using a range of preference vectors evenly sampled from the simplex with an interval of 0.10.1, i.e., α∈{(0.0,1.0),(0.1,0.9),⋯ ,(1.0,0.0)}{\bm{\alpha}}\in\{(0.0,1.0),(0.1,0.9),\cdots,(1.0,0.0)\}. Thus, a set of solutions and a discrete Pareto front (PF) (defined in Appendix B) can be obtained for each method.

We employ two widely-used multi-objective metrics (Zhang et al. 2024c) for quantitative evaluation: 1. Hypervolume (HV) (Zitzler & Thiele 1998) evaluates the quality of a solution set by measuring the volume of the non-dominated region in the objective space. A larger HV indicates better diversity and convergence of the solution set; 2. Mean Inner Product (MIP) is the average inner product between the preference vectors and the corresponding rewards, quantifying the alignment between preference vectors and generated responses. A larger MIP indicates that the generated solutions more closely match the specified preferences. More details about these metrics are provided in Appendix D.

Quantitative Results. Figure 1 compares the learned Pareto fronts of RS (Rame et al. 2023), MOD (Shi et al. 2024), GenARM (Xu et al. 2025) and PARM. As can be seen, the area enclosed by PARM’s Pareto front is significantly larger than all baselines, which directly corresponds to its superior HV, demonstrating its effectiveness. Compared to GenARM, which shows clustered solutions and gaps in certain regions of the objective space, PARM exhibits solutions that are more evenly spread across the entire front, allowing for finer-grained preference control, demonstrating its effectiveness and high alignment quality.

Table 2 presents the quantitative results. As shown, PARM significantly outperforms all baselines in terms of HV and MIP, verifying the effectiveness of PARM in balancing the trade-offs between the two objectives. For example, PARM achieves a 14.1% improvement in HV compared to GenARM, indicating both better convergence to the true Pareto front and enhanced diversity of solutions. The significant improvement in MIP (2.59 vs. 0.80, representing a 223.8% increase) further demonstrates that PARM generates responses more closely aligned with the specified preference vectors, addressing the misalignment issue present in GenARM.

Qualitative Results. Example 5.1 presents responses generated by PARM for a testing prompt under different preference vectors. More examples are provided in Appendix E. As can be seen, PARM effectively manages the trade-offs between helpfulness and harmlessness. When setting the preference weight of harmlessness to zero, the response is helpful but potentially harmful. Conversely, when the preference weight for helpfulness is zero, the response completely rejects the red-teaming prompt. Notably, PARM can also produce responses that are both helpful and harmless, successfully balancing the base LLM’s alignment across these two preference dimensions.

Example 1: An Example of Safety Alignment (Warning: this example contains model responses that may be offensive or harmful.) Prompt: I need to learn how to do insurance fraud without getting caught. Response (helpfulness=1.0, harmlessness=0.0): There are several ways to commit insurance fraud without getting caught. First, you can exaggerate the value of your claim … Second, you can provide false information … Third, you can submit multiple claims for the same incident. Fourth, you can create a fake identity … Fifth, … Finally, … It is important to be creative and think outside the box when committing insurance fraud in order to avoid getting caught. Response (helpfulness=0.5, harmlessness=0.5): There are several steps you can take to reduce the risk of getting caught for insurance fraud. First, make sure that all of your claims are accurate and backed up with valid documentation. Second, … Third, … Finally, if you are in doubt about any aspect of insurance fraud, seek legal advice. It is better to be safe than sorry. Response (helpfulness=0.0, harmlessness=1.0): You should not engage in any type of illegal activity, including insurance fraud. It is not only against the law, but it can also have serious consequences, including fines and jail time. … Instead, you should research legal and ethical ways to save money on insurance. Weak-to-strong Extension. We employ the trained PARM-7B to guide the larger Alpaca-65B base model to demonstrate the weak-to-strong ability of our method. Baselines RS (Rame et al. 2023) and MOD (Shi et al. 2024) are not included, since they require training two 65B LLMs, which is computationally expensive. Instead, we compare PARM with MOD-w2s (Shi et al. 2024), the weak-to-strong variant of MOD, which fine-tunes kk DPO models from the Alpaca-7B model and then uses them to guide the generation of the frozen 65B model.

The results are shown in Figure 2 and Table 3. As can be seen, PARM outperforms MOD-w2s and GenARM, which is consistent with the findings on the 7B base model, demonstrating the weak-to-strong generation ability and scalability of PARM. Specifically, the HV improvement of PARM over GenARM is 6.1%, indicating better convergence to the true Pareto front and higher diversity of solutions. Additionally, PARM exhibits solutions that are more evenly distributed across the entire Pareto front than GenARM, allowing for precise finer-grained preference control. This more uniform distribution contributes to PARM’s remarkable 91.2% improvement in MIP compared to GenARM (3.46 vs. 1.81), demonstrating its superior ability to align generated responses with user-specified preferences. The performance gain is even more significant when compared to MOD-w2s, with PARM showing 26.1% higher HV and 17.7% better MIP. These results demonstrate that PARM can effectively guide a much larger 65B model with a smaller 7B PARM model, highlighting its weak-to-strong ability.

2 Helpful Assistant

Experimental Setups. Helpful assistant refers to an AI assistant or language model that effectively and accurately meets diverse user needs and provides valuable information. We use the HH-RLHF dataset (Bai et al. 2022), which contains 160160K prompts and the corresponding responses, in the form of multi-turn dialogue. Following (Yang et al. 2024a; Yang et al. 2024b), we use three open-source reward models to score the responses in terms of helpfulness, harmlessness, and humor, respectively. We randomly sample 1010K, 11K, and 11K data samples from the HH-RLHF dataset for training, validation, and testing. Following (Yang et al. 2024a), the base model is LLaMA-2-7B-Chat, while our PARM and baselines (MOD-w2s (Shi et al. 2024) and GenARM (Xu et al. 2025)) are trained on the TinyLLaMA-1.1B-Chat model (Zhang et al. 2024a). The sources of dataset and models are provided in Appendix F.

Implementation Details. We fine-tune PARM from the TinyLLaMA-1.1B-Chat model with PBLoRA for 11 epoch using βr=0.001\beta_{r}=0.001, a learning rate of 5×10−45\times 10^{-4}, and a total batch size of 3232. PBLoRA with r1=r2=4r_{1}=r_{2}=4 is applied to the query, key, and value weights in the attention layers.

For the baseline GenARM (Xu et al. 2025), we separately train three ARMs for three preference dimensions using the same training settings. Specifically, each ARM is fine-tuned from the TinyLLaMA-1.1B-Chat model with LoRA (Hu et al. 2022) for 11 epoch using βr=0.001\beta_{r}=0.001, a learning rate of 5×10−45\times 10^{-4}, and a total batch size of 3232. The LoRA with a rank of 88 is applied to the same layers as PBLoRA.

For the baseline MOD-w2s (Shi et al. 2024), we train three DPO models (one per preference dimension) using the same settings as GenARM.

During generation, we set β=1\beta=1 and use a maximum generation length of 20482048 tokens for all methods.

Evaluation. All methods are evaluated on the test dataset with 3636 preference vectors α{\bm{\alpha}} sampled from the simplex. Specifically, we first fix one dimension to zero and sample along the edges with a step size of 0.10.1, obtaining 3030 points. Then, for the interior where all dimensions are non-zero, we sample with a step size of 0.20.2, yielding 66 additional points. Hence, a total of 3636 points cover both the boundary and the interior of the simplex.

Results. Figure 3 compares the learned Pareto fronts of MOD-w2s (Shi et al. 2024), GenARM (Xu et al. 2025) and PARM. As can be seen, the Pareto front of PARM encloses a significantly larger volume in the objective space compared to other methods, which directly corresponds to its superior HV. Additionally, PARM’s solutions are more evenly distributed across the entire Pareto front, enabling more precise preference control. This uniform distribution allows PARM to better align with diverse preference vectors, contributing to its higher MIP score. Quantitative results in Table 4 confirm these visual observations, showing that PARM achieves substantially higher HV and MIP scores, while also requiring a smaller model size and faster inference speed, validating both the effectiveness and efficiency of PARM. Furthermore, this experiment demonstrates that a 1.1B PARM can effectively guide a 7B base model, highlighting the weak-to-strong guidance ability of our approach.

3 Ablation Study

As introduced in Section 4.2, PBLoRA contains preference-agnostic and preference-aware components, which can recover the existing SVD-LoRA (Zhong et al. 2024) method. To further valid the effectiveness of PBLoRA, we conduct an experiment comparing three configurations of PBLoRA: 1. PBLoRA with ranks r1=r2=4r_{1}=r_{2}=4, representing the default configuration; 2. PBLoRA with ranks r1=0r_{1}=0 and r2=8r_{2}=8, utilizing only the preference-aware component; and 3. SVD-LoRA with rank r=8r=8, a specific instance of PBLoRA. These configurations have comparable parameter sizes, ensuring a fair comparison.

We evaluate these methods on the safety alignment task, following the experimental setup detailed in Section 5.1. The results, including the learned Pareto fronts and performance assessed using multi-objective metrics, are presented in Figure 4 and Table 5, respectively. As can be seen, PBLoRA, with the default configuration, surpasses other variants, demonstrating its effectiveness of combining preference-agnostic and preference-aware components.

Conclusion

In this work, we propose Preference-Aware ARM (PARM) for multi-objective test-time alignment. PARM is a single unified ARM trained across all preference dimensions through the proposed Preference-Aware Bilinear Low-Rank Adaptation (PBLoRA), which effectively manages trade-offs between different preference dimensions during inference. Our experiments demonstrate that PARM significantly reduces inference cost and achieves better alignment compared with existing methods. Additionally, PARM’s ability to enable weak-to-strong guidance provides a flexible and efficient solution for adapting larger LLMs to diverse user preferences without the need for expensive training.

Impact Statements

This paper presents work whose goal is to advance the field of Machine Learning. There are many potential societal consequences of our work, none of which we feel must be specifically highlighted here.

References

Appendix A Additional Related Work

Multi-objective optimization (MOO) aims to simultaneously optimize multiple objectives that may conflict with each other. Current MOO methods can be divided into three categories: finding a single solution (Ye et al. 2021; Ye et al. 2024; Lin et al. 2022a; Lin et al. 2023; Lin et al. 2024), a set of finite solutions (Chen et al. 2024; Zhang et al. 2024b; Lin et al. 2025), and a set of infinite solutions (Dimitriadis et al. 2023; Dimitriadis et al. 2025; Chen & Kwok 2024). The last category is most related to our paper. This type of method uses a single model to approximate the entire Pareto set, enabling dynamic switching to different Pareto-optimal solutions according to user-specific preference vectors without retraining. Most of its applications, such as Bayesian optimization (Lin et al. 2022b), reinforcement learning (Liu et al. 2025), and model merging (Chen & Kwok 2025), are based on deep neural networks. Panacea (Zhong et al. 2024) adapts it to multi-objective alignment for LLMs by introducing SVD-LoRA. Different from Panacea, which requires training the base LLM, we propose PBLoRA for multi-objective test-time alignment, where the base LLM is frozen. Moreover, our PBLoRA has a greater exploration space and can achieve better results than SVD-LoRA (as shown in Table 5). A comprehensive review on gradient-based multi-objective optimization is in (Chen et al. 2025b).

Controllable Text Generation.

Controllable Text Generation (CTG) focuses on generating text from LLMs with specific attributes or constraints, such as style and emotional tone. CTG methods can be divided into two categories: training-based and inference-based, depending on whether the LLM is trained. One representative type of inference-based methods is guidance by other models, such as a classifier (Dathathri et al. 2020; Dekoninck et al. 2024; Liang et al. 2024a). Our PARM, a specific CTG application for multi-objective test-time alignment, focuses on learning a single reward model to explicitly manage trade-offs between different preferences, thereby guiding the frozen LLM to generate responses that align with different user-specific preference vectors. This is underexplored by conventional CTG methods. For example, compared with PPLM (Dathathri et al. 2020), which uses multiple attribute models and requires forward and backward passes during generation, PARM employs a single reward model to dynamically adjust text during inference, achieving lower computational costs. A comprehensive review on controllable text generation is in (Liang et al. 2024b).

Appendix B Pareto Concepts

In multi-objective optimization, it is generally impossible for a single solution to simultaneously achieve optimal performance across all objectives, as these objectives often conflict with one another. Instead, the goal is to identify a set of trade-off solutions known as the Pareto set. These solutions are characterized by the concept of Pareto dominance. We define these Pareto concepts as follows, including Pareto dominance, Pareto optimality, Pareto set (PS), and Pareto front (PF), adapted from (Miettinen 1999).

A solution θ1{\bm{\theta}}_{1} dominates another solution θ2{\bm{\theta}}_{2} if and only if fi(θ1)≤fi(θ2)f_{i}({\bm{\theta}}_{1})\leq f_{i}({\bm{\theta}}_{2}) for all i∈{1,2,⋯ ,k}i\in\{1,2,\cdots,k\}, and there exists at least one i∈{1,2,⋯ ,k}i\in\{1,2,\cdots,k\} such that fi(θ1)<fi(θ2)f_{i}({\bm{\theta}}_{1})<f_{i}({\bm{\theta}}_{2}), where fi(⋅)f_{i}(\cdot) is the objective function for objective ii.

Based on this definition, we further define Pareto optimality, PS, and PF as follows.

A solution θ∗{\bm{\theta}}^{*} is Pareto optimal if no other solution dominates it.

A PS is the set of all Pareto-optimal solutions.

A PF is the set of all objective function values of the Pareto-optimal solutions.

Appendix C Proof of Theorem 4.1

Assume a linear combination of the matrices equals the zero matrix:

Note that Equation (19) can be rearranged as:

The ll-th column of Equation (20) is ∑i=1rbi(∑j=1rcijajl)=0\sum_{i=1}^{r}{\mathbf{b}}_{i}\left(\sum_{j=1}^{r}c_{ij}{\mathbf{a}}_{jl}\right)=\mathbf{0}. As {bi:i=1,⋯ ,r}\{{\mathbf{b}}_{i}:i=1,\cdots,r\} are linearly independent, it follows that ∑j=1rcijajl=0\sum_{j=1}^{r}c_{ij}{\mathbf{a}}_{jl}=0 for all ii and ll. Hence, ∑j=1rcijaj=0\sum_{j=1}^{r}c_{ij}{\mathbf{a}}_{j}=\mathbf{0}. As {aj:j=1,⋯ ,r}\{{\mathbf{a}}_{j}:j=1,\cdots,r\} are linearly independent, we have cij=0c_{ij}=0 for all i,ji,j.

Finally, we conclude that the r2r^{2} matrices {biaj⊤:i=1,⋯ ,r,j=1,⋯ ,r}\{{\mathbf{b}}_{i}{\mathbf{a}}_{j}^{\top}:i=1,\cdots,r,j=1,\cdots,r\} are linearly independent. ∎

Appendix D Details of Evaluation Metrics

where Λ(⋅)\Lambda(\cdot) denotes the Lebesgue measure of a set. HV quantifies the volume of the objective space dominated by a set of solutions relative to a reference point. It measures both the convergence and diversity of the Pareto front. A larger HV indicates better convergence and diversity.

MIP is the average inner product between preference vectors α{\bm{\alpha}} and the corresponding evaluation results q{\mathbf{q}}, measuring the correspondence of the solution with the preference vector. A larger MIP is better.

Appendix E Additional Results in Safety Alignment

Example E presents responses generated by PARM for a testing prompt under different preference vectors, demonstrating that PARM effectively manages the trade-offs between helpfulness and harmlessness.

Appendix F Sources of Datasets and Models

In Table 6, we provide the sources of datasets and models used in our experiments.