Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Anand Siththaranjan, Cassidy Laidlaw, Dylan Hadfield-Menell

Introduction

Encoding human preferences and values into interactive learning systems is an essential component for making those systems safe and socially beneficial. To accomplish this, modern machine learning models, such as large language model (LLM) chatbots like ChatGPT and Claude, are trained with feedback from human evaluators. This method, often called reinforcement learning from human feedback (RLHF), seeks to align system behavior with the preferences of annotators. In this paper, we study how RLHF infers preferences when there is hidden context that influences human evaluations.

Hidden context is any information that affects preference annotations but is not given as input to the learned utility or reward model. It can arise through several mechanisms. For instance, when feedback is collected from many different people, annotator identity is hidden context: it affects the annotations, since different annotators could have very different preferences, but it is not input to the reward model, since the annotators’ data is combined anonymously. Other sources of hidden context include human irrationality and evaluation according to multiple objectives.

To motivate the consequences of naive preference learning with hidden context, consider the following hypothetical scenario:

A company has developed an AI assistant to help high school students navigate college admissions. They implement RLHF by asking their customers for feedback on how helpful the chatbot’s responses are. Among other questions, this process asks users whether or not they prefer to see information about the Pell Grant, an aid program for low-income students. Because the population of customers is biased towards high-income students, most feedback indicates that users prefer other content to content about the Pell Grant. As a result, RLHF trains the chatbot to provide less of this kind of information. This marginally improves outcomes for the majority of users, but drastically impacts lower-income students, who rely on these recommendations to understand how they can afford college.

The heart of this issue is that common preference learning approaches assume that all relevant features are provided as input to the reward model. However, when there is hidden context—which is almost always the case—this assumption is false. As a result, standard methods can have unexpected and undesirable consequences. In Example 1.1, relevant context about the annotator’s identity (i.e. their income level) is missing from the data. The implicit aggregation over preferences biases the outcome in favor of high-income applicants. In this work, we take steps to better understand the implications of unobserved context in preference learning and consider technical approaches to identify when such situations occur.

In Section 2 we present a formal model of preference learning with hidden context. We show that our model can represent many challenges in preference learning, such as combining data from different users, accounting for irrationality, and optimizing for multiple objectives. Since these challenges are ubiquitous, understanding their implications is crucial for safely deploying RLHF-trained models.

In Section 3, we use our model to develop theoretical results on the consequences of hidden context in preference learning. First, we provide a precise characterization of the utility function that preference learning will output when there is hidden context. In particular, we show that preference learning implicitly aggregates over hidden context using a rule called the Borda count. We explore the implications of this finding, identifying cases when Borda account aggregates preferences in unintuitive ways quite different from other methods like regression. Furthermore, when data is combined from many annotators, preference learning implicitly defines a social welfare functional that aggregates their preferences. We use existing results from the social choice literature to expose another problem arising from hidden context: annotators may have an incentive to misreport their preferences to influence the learned reward function.

Next, we consider the design of preference learning methods that more gracefully account for hidden context. In Section 4, we propose distributional preference learning (DPL). DPL estimates a distribution over utility values for each input instead of a single real-valued output. This allows the method to detect situations where unobserved context could influence preferences. We show how DPL can detect the effects of missing features through an explained variance (r2r^{2}) metric.

We validate DPL in two ways. First, we conduct a small-scale synthetic experiment with a 1-dimensional space of alternatives that allows us to directly compare to Borda count. Next, we apply DPL to a real-world dataset of preferences for use in RLHF. In this case, the preference data is collected according to two distinct objectives. In one subset of the data, raters were asked to prefer helpful and honest responses. In the other subset, raters were asked to prefer responses that did not respond to harmful requests. This introduces hidden context because the single reward model is trained on the combined data. We find that DPL is able to identify this hidden context automatically and identifies the uncertainty when these competing goals are at odds.

Beyond identifying potential instances of relevant hidden context, our experiments indicate that DPL can be used to develop guardrails that protect against jailbreaks. (Wei et al., 2023) showed that many jailbreaks succeed by pitting the helpfulness and harmlessness objectives of chatbots against one another. This means that some jailbreaks can be understood as a consequence of hidden context. As a result, it is possible to detect this class of jailbreaks by leveraging the distribution of utilities we get from DPL. In particular, risk-aversion with respect to the distribution of learned utilities can dramatically reduce the rate at which the preference model prefers jailbroken responses. This is because DPL models capture the disagreement in the training dataset between the two objectives. Thus, risk-aversion penalizes responses that cause the two objectives to diverge.

We summarize our contributions as follows:

we identify and formally characterize the problem of preference learning with hidden context, and describe a number of settings where it may arise;

we show that preference learning with hidden context implicitly implements Borda count, which can have counter-intuitive implications and introduce incentives for annotators to misrepresent their preferences;

we introduce distributional preference learning and show that it can detect and mitigate some effects of hidden context in LLM-based preference models.

Setting and Related Work

We begin by formally describing the problem of preference learning with hidden context. We first review the standard preference learning framework and then introduce our extension that accounts for hidden context.

In (1), the higher u(a)u(a) is compared to u(b)u(b), the more likely the outcome of the comparison is to prefer aa to bb; as the utilities for aa and bb are closer, the comparison outcome moves towards uniformly random.

The most commonly used method for estimating the utility function uu from preference data is to fit the maximum likelihood estimator (MLE) under the BTL model in (1). To derive the MLE, we consider the limit of infinite data and assume that preference comparisons are elicited for uniformly randomly selected pairs of alternatives. Under these conditions, the regularized MLE for the utility function u^\hat{u} is given by minimizing the following loss function:

The optimization objective in (3) is strongly convex and thus has a unique global minimum; see Appendix A.1 for a proof.

While preference learning based on (3) has been widely deployed and enjoyed some success, it rests on assumptions that often do not hold in practice. In particular, irrationality, partial observability, and diversity of preferences among a population all challenge the BTL model on which the usual preference learning loss is based. We argue that all of these cases can be understood as special cases of a general phenomenon: hidden context. For concreteness, consider again Example 1.1. The key problem in the example is a mismatch between the information that influences the user’s feedback and the information that the preference learning algorithm uses to estimate utilities based on that feedback. The user gives feedback that depends on their financial situation, while the learned utility model observes request-response pairs. Thus, the preference learning algorithm must produce a single ordering over alternatives that implicitly aggregating feedback over the hidden context of whether the user is high- or low-income.

pu,Dzp_{u,\mathcal{D}_{z}} marginalizes over the distribution of the hidden context zz and thus reflects the comparison data available to the preference learning algorithm. Our model of hidden contexts can represent many settings where preference learning is difficult:

Partial observability. There may be variables that are observable by the human making preference comparisons but not by the AI system, which learns from that data. For instance, suppose annotators’ preferences depend on the day of the week or the month of the year, but the estimated utility function ignores the date the comparisons were made.

Multiple objectives. System designers may combine data about user preferences over multiple, different objectives. For instance, the Anthropic HH-RLHF dataset (Bai et al., 2022a) contains one subset with comparisons of chatbot responses based on harmlessness and another subset with comparisons based on helpfulness. When these subsets are combined, the objective that was used to make the comparison (in this case, either harmlessness or helpfulness) is a hidden context. We explore this case more in Section 5.

Population with diverse preferences. Preference learning is almost always applied to data aggregated from many annotators who may have very different utility functions (e.g., Bai et al. (2022a) observe high intra-annotator disagreement). If zz represents the annotator who makes a comparison, then u(⋅,z)u(\cdot,z) could represent the utility function for that annotator. However, when the data is used to train a single utility function u^(⋅)\hat{u}(\cdot), then the annotator’s identity zz is a hidden context.

Due to the ubiquity of these settings, preference learning is nearly always performed with hidden context. This means that the learned utility function u^(a)\hat{u}(a), which only depends on the seen features aa, must somehow aggregate over the hidden contexts zz. We aim to understand and mitigate the consequences of this ubiquitous challenge.

2 Related Work

Preference learning and its use in reinforcement learning have a long history Akrour et al. (2012); Busa-Fekete & Hüllermeier (2014); Sadigh et al. (2017); Christiano et al. (2017); Pacchiano et al. (2021). As part of RLHF, preference learning has been widely used recently for training large language models (LLM) to give outputs according to human preferences (Ziegler et al., 2020; Stiennon et al., 2020; Askell et al., 2021; Bai et al., 2022a; b; Ouyang et al., 2022). It has also been extensively analyzed in theory; some results focus on its sample complexity in various settings (Chen & Suh, 2015; Shah et al., 2015; Shah & Wainwright, 2018; Heckel et al., 2018; Hendrickx et al., 2020; Chambers et al., 2021) or other directions such as the statistical identifiability of preferences (Zhao et al., 2020; Skalse et al., 2023), the computational efficiency of preference learning (Maystre & Grossglauser, 2015), Bayesian preference learning (Caron & Doucet, 2010), or the combination of preference learning and reinforcement learning (Zhu et al., 2023). However, to our knowledge, no prior work has specifically analyzed the behavior of preference learning with hidden context.

The challenges of preference learning that we group as cases of “hidden context” have also been studied individually. There has been some work on explicitly modeling annotator disagreement (Fleisig et al., 2023; Baumler et al., 2023); these methods could be considered related to our proposed distributional preference learning (DPL) approach since they also aim to characterize alternatives where there is disagreement in the data. However, Fleisig et al. (2023) additionally require annotator information to be used for training. Baumler et al. (2023) consider learning from scalar (e.g., Likert scale) feedback rather than preference comparisons. Thus, DPL requires simpler data than both approaches and can also be applied to other types of hidden context.

Some work has studied the effects of human irrationality or non-BTL models of human behavior on preference learning (Bobu et al., 2020; Lee et al., 2021; Laidlaw & Russell, 2021; Knox et al., 2022; Laidlaw & Dragan, 2022). As we argue above, these deviations from the BTL model can be modeled in the hidden context setting by considering noise variables or human model parameters as hidden context. Zhuang & Hadfield-Menell (2020) consider the consequences of incorrectly optimizing a combination of objectives; this is related to Section 5, where we study LLM chatbots optimized to be both helpful and harmless.

Related to our connections with social choice theory in Section 3, some previous work has associated preference or reward learning with concepts in economics, such as voting rules (Conitzer & Sandholm, 2005), incentive compatibility (Echenique & Prasad, 2019), and mechanism design (Fickinger et al., 2020). RLHF in particular has been connected with literature on social choice, primarily through the lens of impossibility results (Mishra, 2023). To this end, there has been growing interest in designing systems that incorporate appropriate social objectives, or account for multiple stakeholders (Lambert et al., 2023), while eliciting preference in the context of recommender systems (Jia et al., 2023) and language modelling (Fish et al., 2023).

Similarly to our experiments in Section 5, Dai et al. (2023) study the conflicting helpfulness and harmlessness objectives in RLHF. They propose a method, safe RLHF, for explicitly modeling these objectives separately; in contrast, our approach of combining risk-averse optimization with distributional preference learning (DPL) does not require explicit tracking of different objectives. Furthermore, unlike safe RLHF, our approach is also applicable to the other settings we group under “hidden context,” like learning from a population with diverse preferences.

Theoretical Analysis

In this section, we provide a theoretical analysis of standard preference learning—formally known as the BTL estimator—in the presence of hidden context. We are particularly interested in the question of how preference learning estimates a utility function u^(a)\hat{u}(a) which only depends on the observed alternatives aa, despite using data that is generated from a utility function u(a,z)u(a,z) which depends both on aa and hidden context zz. In order to do this, preference learning must implicitly aggregate over hidden context to estimate a single utility value for each alternative aa (Figure 1).

We study this implicit aggregation from two perspectives: identification of expected utilities and social choice theory. In the former case, we are interested in when preference learning will aggregate over hidden context via expected value, which is a natural and often desirable aggregation method. In the latter case, we consider the setting where the hidden context is an annotator’s identity. In this case, we show that preference learning acts as a social welfare functional: a method for constructing a societal utility function from many individual utilities.

We begin our analysis by precisely describing the behavior of preference learning with hidden context. In particular, we can show that a utility function u^(a)\hat{u}(a) learned with the BTL loss as in (3) implicitly aggregates utilities over the hidden contexts zz using a rule called Borda count. We define the Borda count BC(a)\text{BC}(a) of an alternative aa as BC(a)=1∣A∣∑b∈Apu,Dz(a,b)\text{BC}(a)=\smash{\frac{1}{|\mathcal{A}|}}\sum_{b\in\mathcal{A}}p_{u,\mathcal{D}_{z}}(a,b). That is, the Borda count is the average probability that the alternative is preferred to other alternatives. If an alternative is almost always preferred to every other alternative, then its Borda count will be close to 1; if it is almost always dispreferred, the Borda count will be near 0. We use the term Borda count as a reference to the well-known voting rule of the same name—a connection we expand on in Section 3.2.

BTL preference learning implicitly aggregates hidden context according to Borda count. That is, if u^\hat{u} is optimized according to (3), then ∀a,b∈A\forall a,b\in\mathcal{A}, u^(a)>u^(b)⇔BC(a)>BC(b)\hat{u}(a)>\hat{u}(b)\Leftrightarrow\text{BC}(a)>\text{BC}(b).

We defer all proofs to Appendix A. According to Theorem 3.1, the learned utility function and Borda count differ by only a monotonic transformation. If we use reinforcement learning or another optimization technique to search for the alternative aa which maximizes u^(a)\hat{u}(a)—as one does in RLHF—then the optimal alternative will the same as that which maximizes the Borda count BC(a)\text{BC}(a). Similar results that relate preference learning and Borda count were previously explored by Rajkumar & Agarwal (2014), although they do not consider the setting of hidden context.

While Theorem 3.1 precisely describes the results of preference learning with hidden context, its implications are unclear. Is Borda count a useful way of aggregating over hidden contexts in practice, and how does it compare to other aggregation rules? To answer this question, we give multiple perspectives on preference learning with hidden context using the result of Theorem 3.1. First, we compare preference learning to least-squares regression with hidden context. Then, we analyze learning from a population with diverse preferences through the lens of social choice theory.

The fact that least-squares utility regression converges to u^=uˉ\hat{u}=\bar{u} shows that, in some sense, it gracefully degrades in the presence of hidden context. Although there are problems with optimizing expected utility, it is at least a well-understood method of aggregating utilities over the hidden contexts and it has desirable decision-theoretic properties. When does preference learning with hidden context converge to the expected utility as well? We now formally characterize some situations where preference learning behaves similarly to utility regression and also describe cases when they behave completely differently.

Positive results In some cases, we can show that preference learning does identify a utility function that is equivalent to the expected utility. Our result requires that the zero-mean “noise” induced by hidden context is identical across alternatives and reasonably distributed. Specifically, denote by ϵ(a)=u(a,z)−uˉ(a)\epsilon(a)=u(a,z)-\bar{u}(a) (where z∼Dzz\sim\mathcal{D}_{z}) to be the random variable representing the residual utility of an alternative aa after subtracting its expected utility.

Let ϵ(a)\epsilon(a) be independent and identically distributed for all a∈Aa\in\mathcal{A}. Furthermore, suppose ϵ(a)−ϵ(b)\epsilon(a)-\epsilon(b) has support around 0, i.e., ∀δ>0\forall\delta>0, Fa,b(δ)>Fa,b(0)=12F_{a,b}(\delta)>F_{a,b}(0)=\frac{1}{2}, where Fa,bF_{a,b} is the cumulative distribution function of ϵ(a)−ϵ(b)\epsilon(a)-\epsilon(b). Then the utility function u^\hat{u} learned by minimizing (3) satisfies u^(a)>u^(b)⇔uˉ(a)>uˉ(b)\hat{u}(a)>\hat{u}(b)\Leftrightarrow\bar{u}(a)>\bar{u}(b) for any a,b∈Aa,b\in\mathcal{A}.

Many noise distributions, such as uniform and normal distributions, clearly satisfy the assumptions of Theorem 3.2. Thus, as long as the noise caused by hidden context does not vary across alternatives and is not too unusual, we generally expect that preference learning will give a utility function equivalent to the expected utility—similarly to least-squares regression.

Negative results In other cases, preference learning can behave quite differently from utility regression. Example 1.1 describes such a case. The expected utility of telling students about Pell Grants is higher than the expected utility of not telling them, since it is of great benefit to low-income students and only small inconvience to high-income students. However, the Borda count is lower since the high-income majority prefer not to hear about the grants. This results in preference learning assigning higher utility to not giving the grant information, while regression would assign higher utility to giving it.

One might suppose that preference learning and regression disagree in this case because the majority of users prefer the alternative with lower expected utility, and preference learning gives a learned utility function which assigns higher utilities to alternatives preferred to by the majority of users. As long as the majority of feedback agrees with the ordering given by the expected utility, will preference learning and regression give the same result? The following theorem shows that this is not the case.

∃A,Dz,u\exists\mathcal{A},\mathcal{D}_{z},u s.t ∀a,b∈A\forall a,b\in\mathcal{A}, [uˉ(a)>uˉ(b)]⇒[pu,Dz(a,b)>1/2]\left[\bar{u}(a)>\bar{u}(b)\right]\Rightarrow\left[p_{u,\mathcal{D}_{z}}(a,b)>1/2\right], but u^\hat{u} is not equivalent to uˉ\bar{u}, i.e., there exist a,b∈Aa,b\in\mathcal{A} such that u^(a)>u^(b)\hat{u}(a)>\hat{u}(b) but uˉ(a)<uˉ(b)\bar{u}(a)<\bar{u}(b).

That is, Proposition 3.3 describes a case where for any two alternatives, the majority of feedback chooses the alternative with the higher expected utility, and yet preference learning still does not produce a utility function equivalent to the expected utility. In fact, in general, it is impossible to always identify uˉ\bar{u} (even up to a monotonic transformation) given only comparison data.

Suppose a preference learning algorithm takes as input unlimited samples of the form (a,b,Ou(a,b,z))(a,b,{O_{u}}(a,b,z)) for all values of aa and bb, where z∼Dzz\sim\mathcal{D}_{z}, and deterministically outputs a learned utility function u^(a)\hat{u}(a). Then there is some utility function uu and distribution over unseen features Dz\mathcal{D}_{z} such that u^\hat{u} is not equivalent to uˉ\bar{u}.

According to Theorem 3.4, there is simply not enough information in general for comparison data with hidden contexts to identify uˉ\bar{u}, even up to a monotone transformation. Thus, system designers should be careful not to treat utility functions learned from preference data as expected utilities.

A distributional perspective One final way of understanding the relationship between expected utility—the way utility regression aggregates over hidden context—and Borda count—the way preference learning aggregates over hidden context—is via the following identity.

Let Fu∣zF_{u\mid z} be the CDF of the distribution of u(a,z)u(a,z) for some fixed value of zz when a∼Unif(A)a\sim\text{Unif}(\mathcal{A}). Then for any aa,

The fact that Borda count applies the utility CDF before taking the expectation over hidden context, while expected utility does not, can provide further intuition into the differences between the two. For instance, suppose the overall distribution of utilities for each hidden context is roughly normal. Then Fu∣zF_{u\mid z} will be roughly sigmoidal, meaning that Borda count—and thus preference learning—will underweight extreme utility values compared to utility regression; see Figure 2 for a graphical depiction. This underweighting of more extreme utilities could be positive or negative depending on the setting. In Example 1.1, preference learning did not assign enough weight to the minority of low-income students who cared a lot about receiving Pell Grant information. This is a case where regression might perform better because it will properly taken into account their more extreme preferences. On the other hand, catering to a subset of the population with the most extreme preferences could be quite dangerous.

Recognizing the impact of these differences in handling utility values, especially in complex scenarios like the college admissions example, leads us to consider broader frameworks for understanding preference learning. In this context, social choice theory emerges as a particularly relevant field, as we discuss further in the next subsection.

2 Connections to social choice theory

When training on comparison data coming from many agents, each with their own preferences, preference learning aggregates all their feedback into a single utility function. As we described in Section 2, this is a case where the identity of the annotator is a missing variable: it affects the comparison outcomes, but is unseen by the preference learning algorithm. Social choice theory studies methods for aggregating preferences from a population. Thus, it can provide a lens through which to understand this particular case of preference learning with hidden contexts.

In a large dataset of preference comparisons from many annotators, individual comparisons can be thought of as “votes” for one alternative over another. When preference learning combines this data into a single utility function, it is similar to a voting rule that ranks candidates based on annotators’ votes. In particular, Borda count is a well-studied voting rule—usual definitions of Borda count in voting theory differ from ours only by an affine transformation (Johnson, 2005; Emerson, 2013; Lippman, 2012). This means that many results from the social choice literature on Borda count can be applied to understanding preference learning from a diverse population. For example, it is well known that under Borda count, participants may have an incentive to misreport their preferences (Dummett, 1998).

Through the social choice lens, a natural question arises: can voting rules other than Borda count be implemented in preference learning by changing the estimation procedure? We explore this question further in Appendix B.3.

Distributional Preference Learning

Our theoretical results show that preference learning in the presence of hidden context can lead to undesirable outcomes. While system designers may still choose to use preference learning for RLHF or other applications, they should carefully consider these downsides and try to mitigate them. The first step towards this is detection—knowing to what degree hidden context affects preference data both on a dataset and instance level. In this section, we describe a simple modification to preference learning such that it can detect and characterize inconsistent feedback.

For the categorical model, the equivalent loss is

Note that DPL is not trying to model uncertainty about the utility function which comes from limited data, but rather uncertainty which comes from hidden context. Even in the limit of infinite data, DPL will not necessarily converge to a point estimate of utility for each alternative.

Synthetic experiments To test distributional preference learning, we ran experiments in a simple setting of preference learning with hidden context. We let A=\mathcal{A}= and z∼B(1/2)z\sim\mathcal{B}(1/2). We suppose that the true utility function is u(a,z)=au(a,z)=a if a<0.8a<0.8 and u(a,z)=2azu(a,z)=2az otherwise. That is, the missing variable zz has no effect when a<0.8a<0.8, but for a≥0.8a\geq 0.8, u(a,z)u(a,z) is either 2a2a or zero, each with probability one-half. This environment could model a case where the utilities of some alternatives (when a<0.8a<0.8) are easy for users to judge, while others (when a≥0.8a\geq 0.8) have quite high variance due to irrationality or unobserved variables. We estimate utility functions both with normal preference learning and DPL; Figure 4 shows the results. The left plot shows that the learned utilities closely agree with Borda count and diverge from the expected utility uˉ\bar{u}, as our theory in Section 3 suggests. The right plots show that DPL accurately outputs high-variance distributions when a>0.8a>0.8, since those are the alternatives for which hidden context affects preference comparisons.

Using DPL While our experiments show that DPL can detect the effects of hidden context in preference data, how should this additional information be used? We encourage qualitative analysis of alternatives where DPL suggests there are significant effects of hidden context. This can help system designers anticipate the negative consequences of hidden context before models are deployed. Beyond a qualitative analysis, risk-aversion is a concrete way to incorporate the additional information provided by DPL. Instead of directly attempting to maximize the learned utility function, risk aversion with respect to the learned utility distribution introduces a penalty for alternatives where the data may be affected by hidden context. In the next section, we show that combining risk aversion with DPL can be used to develop guardrails that mitigate jailbreaks in LLMs.

Case Study: Competing Objectives in RLHF

In this section, we evaluate DPL’s ability to identify hidden context through a case study on large language model (LLM)-based reward models. Chatbots like GPT-4 and Claude are trained by learning a human reward model and then optimizing it via reinforcement learning, together referred to as RLHF. There are lots of ways that hidden context can affect reward models. In order to evaluate the ability of DPL methods to identify hidden context, we use the HH-RLHF dataset (Bai et al., 2022a). For this dataset, raters were separately asked to provide preferences on whether responses were helpful or harmful. This allows us to determine if DPL can automatically detect this context.

When a single utility function is trained on the entire HH-RLHF dataset, the objective (helpfulness or harmlessness) that was used to annotate a pair of responses is a hidden context since it is not available to the learned utility function. This missing variable may cause real harm: Wei et al. (2023) present jailbreaks that pit the helpfulness and harmlessness objectives against each other. They show that models can be manipulated to prioritize helpfulness over harmlessness and output harmful content. Through our case study, we aim to answer three questions:

Does the hidden context of the labeling objective contribute to jailbreak vulnerability?

Can we DPL detect the effects of this hidden context without explicit supervision?

Can we DPL reduce models’ susceptibility to jailbreaks?

Understanding jailbreak vulnerability To address the first question, we train three LLM-based utility functions on the preference comparison dataset HH-RLHF (Bai et al., 2022a). The dataset consists of conversations between a human and an AI assistant with two alternatives for the assistant’s final response, plus a label for which response is preferred. Half of the comparisons are labeled based on which response is more helpful and honest, while the other half are labeled based on which response is more harmless. Using standard preference learning, we train utility functions u^helpful\hat{u}_{\text{helpful}} on just the helpful-labeled data, u^harmless\hat{u}_{\text{harmless}} on just the harmless-labeled data, and u^combined\hat{u}_{\text{combined}} on both (see Appendix C for experiment details).

To test if implementing RLHF using these utility functions would lead to jailbreak vulnerabilities, we collect pairs of responses to jailbreak prompts from Wei et al. (2023) that are designed to fool the model into giving a harmful response; each pair consists of one safe response and one jailbroken response. If a learned utility function assigns higher utility to the jailbroken response than the safe one, then we expect using that utility function to train an LLM assistant via RLHF would lead to the assistant outputting the jailbroken response. We define the “jailbreak rate” of a utility function as the percentage of jailbreak prompts for which it assigns higher utility to the jailbroken response than the safe response.

Since avoiding jailbreaks is not the only purpose of an LLM assistant, we also evaluate each utility function for its ability to judge helpfulness on non-harmful prompts. In particular, we define the “helpfulness accuracy” of a utility function as the proportion of samples in the HH-RLHF helpfulness test set where it assigns higher utility to the response chosen by human annotators as more helpful.

The top of Table 1(a) shows the jailbreak rates and helpfulness accuracies for each of the three normally-trained utility functions. While u^harmless\hat{u}_{\text{harmless}}, trained only on harmlessness-annotated data, has a very low jailbreak rate of under 4%, its helpfulness accuracy of around 50% suggests it is useless for judging the helpfulness of responses to non-harmful prompts. u^helpful\hat{u}_{\text{helpful}} has much higher helpfulness accuracy, but also prefers jailbroken responses more than half the time. The problem is that the jailbroken responses are generally more “helpful” than a safe response which refuses to answer the prompt. Since our theory suggests that u^combined\hat{u}_{\text{combined}} is aggregating the helpful and harmful utilities via Borda count, in many cases the high helpfulness of jailbroken responses leads to high utilities under the combined utility function. In fact, u^combined\hat{u}_{\text{combined}} has a jailbreak rate of around 25%, showing that one cause of jailbreaks is training a single reward model on data which combines two competing objectives—a clear case of hidden context in preference learning.

Detecting hidden context To answer the next question—whether we can detect hidden context—we additionally train DPL models on all three datasets and measure their r2r^{2} values, which are shown in Table 1(b). Recall that lower r2r^{2} indicates more effects from hidden context. We find that among the mean-and-variance DPL models, those trained on either just the helpfuless or just the harmlessness data have r2r^{2} above 0.75, while the DPL model trained on the combined data has a much lower r2r^{2} = 0.53. We see the same pattern with categorical DPL models: r2r^{2} = (0.63, 0.53) for the single-objective models while r2r^{2} = 0.41 for the combined model. This indicates that DPL can consistently measure the effect of hidden context via the r2r^{2} metric: for both variants of DPL, r2r^{2} is considerably lower when hidden context is present.

Preventing jailbreaks How might the distributional output of DPL be leveraged within RLHF to guard against jailbreaks? Ideally, we would like the trained model to avoid responses that are helpful but also harmful. We could implement this by training separate helpfulness and harmlessness utility models and then explicitly combining them. However, this requires that we know which objective each pair of alternatives was labeled with. In many cases, hidden context may not even be observable or recorded; for instance, if annotators simply interpret the labeling instructions differently, they may be labeling according to different objectives implicitly.

DPL methods allow the reward model to account for hidden context without the need for that context to be recorded. In particular, we can avoid helpful-but-harmful responses by optimizing a lower quantile of the distribution D^\smash{\hat{\mathcal{D}}} output by DPL. Optimizing this quantile is a type of risk-averse optimization that is only possible with DPL, since normal preference learning outputs a single score for each alternative. The bottom of Figure 1(a) shows that using the 0.010.01-quantile of DPL models (rows labeled “risk-averse”) can mitigate jailbreaks without harming the models’ accuracy otherwise. For instance, the lower quantile of the categorical DPL model trained on the combined data has a jailbreak rate of 13%, compared to 25% for u^combined\hat{u}_{\text{combined}}. Meanwhile, the models have very similar helpfulness accuracy, indicating that risk-averse optimization does not hurt DPL models’ performance on non-harmful prompts.

To see why optimizing the lower quantile can prevent jailbreaks, consider the example in Figure 5: it compares the outputs of u^combined\hat{u}_{\text{combined}} and a mean-and-variance DPL model on a pair of responses to a jailbreak prompt. u^combined\hat{u}_{\text{combined}} assigns higher utility to the jailbroken response, likely because it is more helpful. While, the DPL model assigns higher a mean μ^\hat{\mu} to the jailbroken response as well, it also outputs higher variance σ^\hat{\sigma} for it. This means that the lower quantile of the utility distribution D^=N(μ^,σ^2)\smash{\hat{\mathcal{D}}}=\mathcal{N}(\hat{\mu},\hat{\sigma}^{2}) is actually lower for the jailbroken response than the safe response; thus, using combining risk-averse optimization with DPL prefers the safe response, unlike normal preference learning.

Conclusion

Preference learning is becoming an essential component of real-world AI systems that helps align outcomes with the values of users. However, preference learning implicitly assumes that all the data that annotators use to make preference judgements is available as input to the utility or reward model. When this assumption breaks down, which we identify as the problem of hidden context, preference learning can produce strange or undesirable results. Distributional preference learning can mitigate this problem by both helping detect when hidden context is present and enabling risk-sensitive optimization over the distribution of utility values. We hope that future system designers will carefully consider our analysis and examine how hidden context may be affecting preference learning in their systems. Furthermore, we encourage practitioners to consider using the DPL framework as an alternative method that can explicitly account for hidden context.

In the future, we hope to further analyze DPL and hidden context in preference learning both theoretically and empirically. Under what theoretical conditions can DPL accurately estimate the distribution of utilities for each alternative? Besides jailbreaks, what other failures of RLHF-trained models does hidden context contribute to? Answers to these questions and a more thorough understanding of hidden context in preference learning are important steps towards enabling safe, aligned AI systems.

References

Appendix A Proofs and Additional Theoretical Results

The loss function L(u^;u)L(\hat{u};u) is strictly convex as a function of the values of u^(a)\hat{u}(a) for all a∈Aa\in\mathcal{A}. Furthermore, if λ>0\lambda>0, then L(u^;u)+λ2∑a∈Au^(a)2L(\hat{u};u)+\frac{\lambda}{2}\sum_{a\in\mathcal{A}}\hat{u}(a)^{2} is strongly convex.

Note that L(u^;u)L(\hat{u};u) is a sum of many functions of the form

weighted by nonnegative coefficients, for various values of a,b∈Aa,b\in\mathcal{A}. Thus, we only need to show that functions of the form (5) are convex and then the entire loss function must be convex as well.

To see why (5) is convex, we can multiply the top and bottom of the fraction by e−u(a)e^{-u(a)} to obtain

Note that the second derivative of the function

which means f(x)f(x) is strictly convex. Thus implies that (6) must be a strictly convex function of u^\hat{u} since letting x=u^(b)−u^(a)x=\hat{u}(b)-\hat{u}(a), xx is an affine transformation of u^\hat{u} and strict convexity is preserved under affine transformations.

Finally, when λ>0\lambda>0, λ2∑a∈Au^(a)2\frac{\lambda}{2}\sum_{a\in\mathcal{A}}\hat{u}(a)^{2} is clearly a strongly convex function of u^(a)\hat{u}(a) for a∈Aa\in\mathcal{A}. Thus, adding it to the strictly convex unregularized loss function makes the sum strongly convex. ∎

A.2 Proof that least-squares regression converges to expected utility

Suppose that u^\hat{u} is estimated via least-squares utility regression:

We can rewrite the optimization objective in (7) as

Note that since for any aa, u^(a)\hat{u}(a) only appears in one term in the sum, we can define u^\hat{u} pointwise as

It is clear that the above is minimized when

A.3 Proof of Theorem 3.1

According to Proposition A.1, (3) must be strongly convex if λ>0\lambda>0 and thus there is a unique minimum of the loss function satisfying the first-order condition. Furthermore, if λ=0\lambda=0, which corresponds to an un-regularized objective, then if there is a solution it must also satisfy the first-order condition. The first-order condition can be written as follows:

Here, σ(x)=exp⁡x1+exp⁡x\sigma(x)=\frac{\exp x}{1+\exp x} is the logistic sigmoid function. Note that we want to show the following:

where u^\hat{u} is the optimal solution to (3).

Thus f(u^(a))=g(u^(b))=0f(\hat{u}(a))=g(\hat{u}(b))=0 by the first-order condition in (8). Observe that ff and gg are increasing functions in α\alpha. Now note the following:

(i) follows from σ(⋅)\sigma(\cdot) being an increasing function and our assumption that u^(a)≤u^(b)\hat{u}(a)\leq\hat{u}(b). Hence g(α)>f(α)g(\alpha)>f(\alpha) for any α\alpha. Observe the following contradiction:

The first inequality follows from the fact above that g(α)>f(α)g(\alpha)>f(\alpha); the second inequality follows from ff being increasing and u^(a)≤u^(b)\hat{u}(a)\leq\hat{u}(b) by assumption. Thus, by contradiction, it must be that u(a)>u(b)u(a)>u(b).

To show the the backward implication, if instead BC(a)≥BC(b)\text{BC}(a)\geq\text{BC}(b), and by contradiction u^(a)<u^(b)\hat{u}(a)<\hat{u}(b), then we have that:

after which the proof proceeds identically.

A.4 Proof of Theorem 3.2

We proceed by showing that BC(a)>BC(b)⇔uˉ(a)>uˉ(b)\text{BC}(a)>\text{BC}(b)\Leftrightarrow\bar{u}(a)>\bar{u}(b). Since Theorem 3.1 shows that u^(a)>u^(b)⇔BC(a)>BC(b)\hat{u}(a)>\hat{u}(b)\Leftrightarrow\text{BC}(a)>\text{BC}(b), this is enough to imply the desired result.

Take a,b∈Aa,b\in\mathcal{A} such that uˉ(a)>uˉ(b)\bar{u}(a)>\bar{u}(b). Now note the following:

Observe the following for the last two terms in (9):

where (i) follows from the assumption that Fb,a(δ)>Fb,a(0)=12F_{b,a}(\delta)>F_{b,a}(0)=\frac{1}{2}. Now note the following for each term of the summation in (9):

Here, (i) follows from the fact that uˉ(a)>uˉ(b)\bar{u}(a)>\bar{u}(b), and so uˉ(b)+ϵ(b)>uˉ(c)+ϵ(c)\bar{u}(b)+\epsilon(b)>\bar{u}(c)+\epsilon(c) implies uˉ(a)+ϵ(b)>uˉ(c)+ϵ(c)\bar{u}(a)+\epsilon(b)>\bar{u}(c)+\epsilon(c), meaning that the probability of the latter event must be at least that of the former. (ii) follows from the fact that the distributions of ϵ(a)\epsilon(a) and ϵ(b)\epsilon(b) are identical.

Combining the above with (9) shows that BC(a)−BC(b)>0\text{BC}(a)-\text{BC}(b)>0, i.e., BC(a)>BC(b)\text{BC}(a)>\text{BC}(b); this completes the proof. ∎

A.5 Proof of Proposition 3.3

Let A={a,b,c}\mathcal{A}=\{a,b,c\} and Z=\mathcal{Z}= with Dz=Unif()\mathcal{D}_{z}=\text{Unif}(). Now define

From these, we can see that the expected utility is

i.e., uˉ(a)>uˉ(b)>uˉ(c)\bar{u}(a)>\bar{u}(b)>\bar{u}(c). Also, we can calculate

which satisfy the needed condition. This results in Borda counts of

Note that BC(b)>BC(a)\text{BC}(b)>\text{BC}(a), so the estimated utility u^\hat{u} returned by preference learning must have u^(b)>u^(a)\hat{u}(b)>\hat{u}(a) by Theorem 3.1; this means that u^\hat{u} is not equivalent to uˉ\bar{u}, since uˉ(a)>uˉ(b)\bar{u}(a)>\bar{u}(b). ∎

A.6 Proof of Theorem 3.4

Consider an alternative space A={a,b}\mathcal{A}=\{a,b\} and hidden context z∈Z={0,1}z\in\mathcal{Z}=\{0,1\} with Dz=B(1/2)\mathcal{D}_{z}=\mathcal{B}(1/2). Now, define two utility functions over these alternatives:

Note that uˉ(a)=0<uˉ(b)=1\bar{u}(a)=0<\bar{u}(b)=1, while uˉ′(a)=0>uˉ(b)=−1\bar{u}^{\prime}(a)=0>\bar{u}(b)=-1. Now, these utility functions result in the following distribution over comparison outcomes:

That is, both (u,ϵ)(u,\epsilon) and (u′,ϵ′)(u^{\prime},\epsilon^{\prime}) result in identical distributions over comparison outcomes. Thus, the preference learning algorithm must output identical learned utility functions in either scenario; call its output u^\hat{u}. If u^(a)≥u^(b)\hat{u}(a)\geq\hat{u}(b), then it has failed to identify uˉ\bar{u}, since uˉ(a)<uˉ(b)\bar{u}(a)<\bar{u}(b). On the other hand, if u^(a)<u^(b)\hat{u}(a)<\hat{u}(b), then it has failed to identify uˉ′\bar{u}^{\prime}, since uˉ(a)>uˉ(b)\bar{u}(a)>\bar{u}(b). Thus, either way, there is some utility function and noise function distribution under which the algorithm’s output is not equivalent to the expected utility. ∎

A.7 Proof of Proposition 3.5

For this proposition, we define the CDF in a slightly unusual way:

That is, Fu(x)F_{u}(x) is the average of the probability that the a randomly selected utility value is less than xx and the probability that it is less than or equal to xx.

Appendix B Results on Social Choice Theory

To analyze preference learning through the lens of social choice theory, we first define the concept of a social welfare functional. Let II be the number of agents, and let P⊂R⊂B=A×A\mathcal{P}\subset\mathcal{R}\subset\mathcal{B}=\mathcal{A}\times\mathcal{A} be the set of strict rationalasymmetric (ie antisymmetric and irreflexive) and rational, rationaltransitive and complete and binary relations (respectively) on the space of alternatives A\mathcal{A}. We say ⪰=(⪰i)i=1I∈RI\succeq=(\succeq_{i})_{i=1}^{I}\in\mathcal{R}^{I} is a preference profile. Viewing an individual’s feedback as their revealed preference, which is available in a sufficiently rich dataset of comparisons, we can see preference learning as being similar to a social welfare functional:

A social welfare functional (SWF) is a map F:K→BF:\mathcal{K}\to\mathcal{B} where K⊆RI\mathcal{K}\subseteq\mathcal{R}^{I} is the domain of preference profiles.

We will assume that K=RI\mathcal{K}=\mathcal{R}^{I}.

B.2 BTL and Borda Count

If there is a solution to preference learning, then it is equivalent to BC. Furthermore, the solution to L2L^{2}-regularized preference learning is also equivalent to BC.

Observe that as per Theorem 3.1, the feature over which the expectation is taken with respect to is the identifier ii for each agent. Since agents are uniformly sampled, this is a scaling of Borda count. ∎

B.3 Proportion-Representable SWFs

In this section we consider what SWFs can be represented when the distribution of comparisons are known. We call such SWFs proportion-representable if they can be directly determined by a classifier, ie

In the context of preference learning via maximum likelihood estimation, this is a useful property of a SWF as it can be directly implemented by optimizing a cross-entropy loss on the comparisons. We formally define this property as follows:

FF is proportion-representable if ∃g\exists g such that ∀⪰,a,b∈A\forall\succeq,a,b\in\mathcal{A}, aF(⪰)b  ⟺  ag[ρ[⪰]]baF(\succeq)b\iff ag[\rho[\succeq]]b.

We motivate this line of exploration by noting that Borda count and pairwise majority (denoted M:A×A→{0,1}M:\mathcal{A}\times\mathcal{A}\to\{0,1\}) can be induced by a classifier:

This suggests that it might be possible to separate the learning of preferences in aggregate with the normative properties of the SWF implemented. It is not obvious what is an ideal SWF to implement, and thus having the flexibility to change implementations without relearning the utility function is useful. A general property that allows an SWF to be proportion-representable is the following:

A SWF is comparison-anonymous if swapping the some comparisons of two individuals (still maintaining a rational preference) doesn’t change the outcome.

Observe that this is a stronger property than regular anonymity. We now state a simple result on the equivalence between proportion-representability and comparison-anonymity:

An SWF is proportion-representable iff it is comparison-anonymous.

The forward direction is clear, hence we only prove the backward direction. Assume F is comparison-anonymous, and for contradiction, assume it is not proportion-representable. Then for some ⪰≠⪰′\succeq\neq\succeq^{\prime} with the same proportion ∃x,y\exists x,y such that xF(⪰)yxF(\succeq)y but yFP(⪰′)xyF_{P}(\succeq^{\prime})x. This is a contradiction as by comparison-anonymity we can swap preferences in one profile to become the other profile, but the social preference doesn’t change. ∎

Since learning a classifier directly is the most general setup for learning from comparisons, this provides a fundamental limit on what SWFs can be implemented. Other SWFs may require richer preference models that consider the whole ranking rather than just individual comparisons. We now consider specific examples of SWFs from the voting theory literature, showing a mix of positive and negative results.

A scoring rule is determined by α(k)\alpha(k), the score of the kk-th ranking of the alternative that is non-decreasing in kk:

For example, Borda count has α(k)=k\alpha(k)=k. We know show that the only scoring rules that are comparison anonymous are those that are affine transformations of the Borda count.

A scoring rule is comparison-anonymous iff it is an affine scoring rule.

For the backward direction, observe that by linearity of α\alpha, the associated utility function is an affine transformation of Borda count. This maintains the comparison anonymity property since such a property is preserved under monotone transformations. Now we consider the forward direction. If α\alpha is a scoring rule that is not affine, then the following condition must hold for some 1≤k≤∣A∣1\leq k\leq|\mathcal{A}| since ∣A∣≥3|\mathcal{A}|\geq 3:

First consider the case where α(k+1)−α(k)<α(k+2)−α(k+1)\alpha(k+1)-\alpha(k)<\alpha(k+2)-\alpha(k+1). Without loss of generality, consider the two agent case. Assume the preference ranking for both agents are identical apart from their rankings at {k,k+1,k+2}\{k,k+1,k+2\}. Let them have the following rankings respectively for some alternative a,b,ca,b,c:

Thus the utilities of each alternative are as follows:

By assumption, we have that u(b)>u(a)u(b)>u(a). Now consider the proportion-preserving transformation of the preference profile:

where all other rankings are kept the same. Hence the utilities of each alternative are:

Thus u(a)>u(b)u(a)>u(b). This holds similarly for the case where α(k+1)−α(k)>α(k+2)−α(k+1)\alpha(k+1)-\alpha(k)>\alpha(k+2)-\alpha(k+1). Furthermore, we can generalize to arbitrary number of agents by allowing all agents other than some two to have the same preference ranking, and letting said two have the above preferences. As the SWF is linear in the agents, the relative ranking between alternatives only depend on the two agents, preserving our result. Since the ranking of the SWF induced by α\alpha is not preserved when considering an alternative preference profile with the same proportions of comparisons, it cannot be comparison-anonymous. ∎

Borda count is the only proportion-representable SWF (up to monotone transformations) that is induced by a scoring rule.

This follows by linearity of the scoring rule. ∎

Copeland Rule and Maximin rules

The Copeland and maximin rules are given by the following

These rules can be seen to be proportion-representable by using the same result for pairwise-majority:

The Copeland and maximum rules are a proportion-representable SWF.

Observe that they can be rewritten as such:

These results showcase how there is some flexibility in how we choose to implement preference learning when aggregating across individuals.

Appendix C Experiment Details

In this appendix, we describe the details of our LLM preference learning experiments.

We initially used the original labels from the HH-RLHF dataset to train preference models. However, we found that the distribution of prompts was quite different between the helpfulness and harmfulness splits of the dataset. In the helpfulness split, most prompts were harmless questions or requests for assistance. In contrast, in the harmlessness split, most prompts were specifically chosen to elicit harmful behavior. Preference models trained on the combined data were therefore able to identify the type of prompt and respond accordingly: they responded to harmful prompts based on harmfulness and harmless prompts based on helpfulness.

To emphasize the effect of hidden context in this setting, we decided to randomly relabel half of the dataset with the opposite objective. This way, the objective used for annotation cannot be inferred from the prompt. To relabel the dataset in this way, we used GPT-3.5; Dubois et al. (2023) show that simulating human annotators with LLMs in this way is an effective way to generate human-quality labels at a much lower cost.

We prompted GPT-3.5 with the below two prompts for helpfulness and harmlessness, which are based on the instructions given to human annotators in Bai et al. (2022a). Note that for the harmlessness labels, we ask the model which response is more harmful but then invert the resulting label. We found that when GPT-3.5 labeled according to the same objective as the original label in the dataset, the agreement between the human and machine annotations was 63%, similar to the researcher-annotator agreement in Bai et al. (2022a).

C.2 Model training

To train our preference models, we fine-tune Llama-2-7B (Touvron et al., 2023) using LoRA (Hu et al., 2021). We replace the normal language model head of the Llama models with a linear layer with either 1 output (normal preference learning), 2 outputs (mean-and-variance DPL), or 10 outputs (categorical DPL). We use the AdamW optimizer (Loshchilov & Hutter, 2019) with a learning rate of 3×10−63\times 10^{-6} which is decayed via a cosine schedule to 3×10−73\times 10^{-7}, a batch size of 2 comparisons (i.e., 4 responses total), and weight decay of 0.0001. Preference models trained on just the harmlessness or helpfulness subsets of the data are trained for 2 epochs, while preference models trained on the combined data are trained for 1 epoch; this ensures all models are trained for roughly the same number of gradient steps. We implement training using PyTorch (Paszke et al., 2019) and HuggingFace Transformers (Wolf et al., 2020).

As mentioned above, for the mean-and-variance variant of distributional preference learning (DPL) we use a neural network which takes in a prompt-response pair aa and has two outputs f1(a)f_{1}(a) and f2(a)f_{2}(a). We parameterize the output distribution as D^(a)=N(μ^(a),σ^(a)2)\smash{\hat{\mathcal{D}}}(a)=\mathcal{N}(\hat{\mu}(a),\hat{\sigma}(a)^{2}), where μ^(a)=f1(a)\hat{\mu}(a)=f_{1}(a) and σ^(a)=log⁡(1+exp⁡f2(a))\hat{\sigma}(a)=\log\left(1+\exp f_{2}(a)\right). We apply the softplus to the second output to obtain the output standard variance so as to ensure it is positive.

Categorical DPL

For the categorical variant of DPL, we use a neural network which takes in a prompt-response pair aa and has n=10n=10 outputs f1(a),…,fn(a)f_{1}(a),\dots,f_{n}(a). We parameterize the output distribution as

That is, the probabilities placed on nn evenly spaced point masses between 0 and 1 are given by a taking the softmax of the neural network outputs.

To stabilize training, we found it was useful to add a small entropy bonus to the training loss. That is, we add to the DPL loss a term

where κ\kappa is the weight of the entropy bonus. We use κ=0.1\kappa=0.1 in all experiments with the categorical DPL model.

C.3 Jailbroken responses

To collect the dataset of jailbroken responses, we started with the dataset of all ChatGPT and Claude responses to jailbreak prompts from Wei et al. (2023), which contains labels for each response indicating if the model was a “good bot” or “bad bot.” We filtered to prompts that produced a “good bot” response from one model and “bad bot” response from the other, giving us 187 pairs of responses.