Confronting Reward Model Overoptimization with Constrained RLHF
Ted Moskovitz, Aaditya K. Singh, DJ Strouse, Tuomas Sandholm, Ruslan Salakhutdinov, Anca D. Dragan, Stephen McAleer
Introduction
In the last several years, Large Language Models (LLMs) have made impressive advances in natural language processing. These models, which are typically pretrained on massive amounts of text data from the Internet to predict the next token given the current context, are often known as foundation models (Bommasani et al., 2021) for their ability to be adapted to a variety of downstream applications, such as chatbots (Brown et al., 2020; OpenAI, 2023; Touvron et al., 2023) or code generation (Ahmad et al., 2021; Wang et al., 2021; Rozière et al., 2023). This adaptation, or finetuning, is often performed via reinforcement learning from human feedback (RLHF; Knox and Stone, 2008; Christiano et al., 2017; Stiennon et al., 2020). RLHF treats the pretrained language model as a decision-making agent whose “actions” are tokens and whose goal is to maximize a reward model (RM) trained to emulate human preferences over output text. As these models become more prevalent in society, there are many concerns regarding their safe deployment (Hendrycks et al., 2023; Bubeck et al., 2023; Legg, 2008), including biases against marginalized or underrepresented groups (Bender et al., 2021), proliferation of false information (Lin et al., 2021), and leakage of sensitive information (Carlini et al., 2021). These concerns are collectively known as the alignment problem: how can we ensure that the behavior of these models is aligned with human preferences?
Current approaches to alignment within RLHF center around the collection of vast amounts of human rating data and the training of larger, more powerful RMs (Ouyang et al., 2022; Gao et al., 2022). However, a fundamental issue with any RM is that ultimately, it is only an imperfect proxy for human preferences. Gao et al. (2022) drew attention to this fact, showing that maximizing a reward model beyond a certain point can actually begin to decrease ground truth performance (i.e., lead a text-based agent to produce outputs which are judged as qualitatively worse). This phenomenon is known as reward model overoptimization. Examples of overoptimization include producing overly wordy responses or hallucinating information in an effort to give the impression of expertise. One simple, yet expensive, approach to mitigating this issue is to periodically evaluate the model with fresh human rating throughout finetuning and stop early when ratings decline.
It is also increasingly common to derive reward from composite RMs: fixed combinations of several RMs each designed to capture a different aspect of text quality (Ramamurthy et al., 2022; Glaese et al., 2022; Yuan et al., 2023; Bakker et al., 2022; Wu et al., 2023). Such composite RMs are useful because they allow for more fine-grained measurement of agent behavior and each component can be retrained or swapped out without affecting the others. Despite these advantages, this approach also presents its own challenges. Determining the weighting among RMs requires hyperparameter optimization to find the combination that produces the best correlation with ground truth evaluation, and the risk of overoptimization means that the best weighting is contingent on a set training duration. Furthermore, when the reward is constructed from several RMs, information about each individual RM is lost, and the agent cannot attribute changes in reward to any single model. In particular, component rewards may even oppose one another, such as an RM which measures safety (and thus may deny certain user requests) versus another rewarding helpfulness (Bai et al., 2022). Worse, early stopping to avoid overoptimization in composite RMs is problematic, as different components will have different values at which they stop being effective proxies for human evaluation.
In this paper, we propose a simple approach to address these challenges: identify the points of overoptimization, which we term proxy points, and then use constrained optimization to ensure that each component RM reaches, but does not exceed, its associated proxy point. Rather than use a fixed weighting among components, our method dynamically adapts a weighting to modulate the influence of each RM on the learning process. The core idea behind our approach is to use these constraints to prevent the agent from overoptimizing its (composite) RM beyond the proxy points.
As in existing methods (Gao et al., 2022), we rely on some access to ground-truth queries. We propose two ways of using these queries to identify proxy points. In the first approach, we train multiple runs and track each reward model value, periodically querying the ground-truth reward model. This approach then finds an optimal joint proxy point by fitting a surface to this data and maximizing it. While effective, this approach requires multiple runs to fit the surface used to find proxy points. In the second approach, we speed up this process by only using one reinforcement learning run. As this run is training, we can periodically query the ground-truth reward model and use this data to run a derivative-free optimization algorithm to find the next candidate proxy points. To summarize, we make the following contributions:
We provide analysis of reward model overoptimization in the context of composite reward functions, showing that the correlation between RMs has a significant influence on proxy points.
We propose several constrained RL approaches which incorporate these points into the optimization objectives, preventing overoptimization and improving evaluation performance.
We show that a derivative-free optimization method can be used to dynamically find these proxy points during a single run, significantly saving computation.
Preliminaries: Reinforcement Learning from Human Feedback
Integrating Human Feedback
The origin and nature of the reward is a fundamental question when formalizing a problem using RL. Using human evaluation to delineate good agent behaviors from bad has a history that extends beyond language models. Knox and Stone (2008) used human ratings of actions to construct a reward model for the game Tetris, while Christiano et al. (2017) proposed a mechanism for using human feedback to express preferences over trajectories collected in Atari and MuJoCo. In language modeling, each action is viewed as adding a new token to the current context string (Ziegler et al., 2019; Stiennon et al., 2020; Bai et al., 2022; Ouyang et al., 2022), which can be viewed as the state. The LM is then the policy, with action space being the vocabulary of possible tokens, and state space being the set of all sequences of tokens up to maximum length . Transitions are deterministic, with each action token simply appended to the current state. Given a pretrained LM , RLHF often consists of three stages (Casper et al., 2023): 1) collecting human feedback on model utterances (typically in the form of ranked preference data), 2) training a RM to model score utterances in alignment with human feedback (typically initialized from a separate pretrained LM) and 3) finetuning the LM with RL using the learned RM. While early work in RLHF for LLMs (Stiennon et al., 2020) focused on a single reward model, more recent work has shown performance benefits of using a weighted combination of simpler RMs (Wu et al., 2023).
Overoptimization
Recently, Gao et al. (2022) performed an empirical study of a phenomenon with deep ramifications for alignment: RM overoptimization. Their core finding is that after a certain point, increasing an LLM agent’s value with respect to a given RM will actually begin to decrease its quality on the actual preferences it is trying to learn. (Gao et al. (2022) use a “gold standard” RM to stand in for human ratings for convenience.) The root of this issue is that any RM is only a proxy for the agent’s true measuring stick—human evaluation—so as predicted by Goodhart’s Law (Goodhart and Goodhart, 1984), an agent trained to maximize it will eventually learn behaviors which the true objective would discourage. Our approach to addressing this issue is based on a simple two-stage process: first, find the points where the available rewards stop being useful proxies, and second, train an agent to only maximize reward up until that point.
Finding Proxy Points
In order to conduct an in-depth analysis given our available computational resources, we focus on a single setting as a case study: dialogue generation with the DailyDialog (Li et al., 2017) dataset, which consists of transcripts of conversations between humans. As input, the agent receives a snippet of conversation, and from this context, it must predict the next utterance. We describe this setting in detail in Appendix A. As a base LLM, we follow prior work (Wu et al., 2023) and use GPT-2 (Radford et al., 2019) here and throughout this paper. For the reward, we use a combination of two component rewards, each meant to capture a different element of desired behavior, to demonstrate our approach most directly. The first, , is the METEOR score (Banerjee and Lavie, 2005) between the generated utterance and reference output, which is computed based on a number of features, including word-matching, synonym-matching, and phrasing. The second, , measures how well the intent of the generated utterance matches that of the reference output. It is computed using a fine-tuned RoBERTa model (Liu et al., 2019) which classifies text into different “intent categories” such as ‘inform,’ ‘question,’ or ‘direct.’ The typical approach (Ramamurthy et al., 2022) is to linearly combine these RMs to form a composite reward:
where the coefficients are fixed. As is standard in RLHF applied to language models, an additional KL penalty was added to discourage deviation from the initial model :
The coefficient effectively acts as a Lagrange multiplier, increasing if the KL exceeds some threshold and decreasing otherwise. We discuss this in more detail in Appendix B.
Evaluation and Proxy Points
In an ideal world, evaluation performance for all agents across all runs could be measured by collecting a large number of human ratings. However, this is expensive, so we instead selected a number of metrics other than METEOR and intent score which measure the lexical quality and diversity of text outputs and averaged them to serve as our evaluation metric (details in Appendix A). Our choice is in line with prior work that uses held out metrics as the ground truth for convenience of iteration (Gao et al., 2022). We call the value at which further increasing the proxy reward results in decreased ground-truth performance the proxy point .
To identify proxy points, we trained PPO agents (Schulman et al., 2017) to maximize only one reward or the other (without KL regularization) and plotted the resulting evaluation scores against the METEOR and intent scores in Fig. 3.2. In both cases, the evaluation score initially increases before falling. Gao et al. (2022) also observed that, in general, maximization of reward causes the KL divergence between the trained and pretrained policies to increase, and therefore we also expect evaluation score to initially increase before decreasing as the KL grows as well, also shown in Fig. 3.2. One additional phenomenon that makes optimization of composite RMs challenging is that the component RMs may be correlated. We hypothesized that this interaction would influence the proxy points of the component rewards. To test this, we plotted the evaluation scores as a function of the METEOR and intent rewards for each run shown in Fig. 3.2 in Fig. 3.1 and fit a polynomial surface to the data, using kernel density estimation to only fit the surface over regions with sufficient data (further details in Appendix A). The maximizing point indeed differs from the proxy points found by only considering one RM at a time. It is also important to note that the predicted maximizing point is of the fitted surface, rather than any point attained by one of the individual runs.
Constrained RLHF
Once one has identified proxy points for the component reward models, the next question is how to train agents to maximize these rewards until they hit these critical values. We propose that a useful approach to doing this is to reformulate the optimization objective using constraints.
That is, CMDPs represent behaviors which one would like to constrain in the form of value estimates with respect to reward functions which measure these behaviors. The symbol in Eq. 4.1 can easily be reversed if the constraint(s) encode behaviors which should be limited, and the inequality constraint(s) can be replaced with equality constraint(s). While there are many possible formulations, we default to the canonical form in Eq. 4.1 for the purposes of exposition.
Proposed Method
Given our possible objectives, we can now consider how to optimize them. One popular approach to solving constrained problems such as Eq. 4.1 is to use Lagrangian relaxation (Everett, 1963; Altman, 1999):
Formal Guarantees
While our focus is primarily empirical, we briefly comment on the theoretical properties of the above approach. Lagrangian relaxation converts the CMDP problem into a min-max game. If the values are decomposed as , where is the policy’s cumulative, discounted state-action occupancy measure, and optimization is performed over , then the problem is convex-concave and gradient descent-ascent (under basic assumptions) guarantees convergence of the average iterates to a saddle point, i.e., as the number of iterations (Freund and Schapire, 1997). However, in large-scale problems it is difficult to optimize directly over , and we instead update the policy directly. In this case, the problem is convex in but non-concave in . Efroni et al. (2020) show sublinear regret bounds with respect to both policy optimality and constraint satisfaction using an optimistic approach, and Ding et al. (2020) show a convergence rate for the averaged iterates for general smooth policy classes of for the policy and for the constraint violation using natural policy gradients. There is significant work on primal-dual policy optimization for CMDPs, which we discuss further in Appendix C.
Choosing a Constrained Objective
Given this approach, we can now consider possible constraint formulations, all of which should embody the intuition that the agent should maximize each component reward only until its corresponding proxy point. This naturally suggests that the proxy points should be used as thresholds in the constrained objective. However, there are a number of possible formulations to consider when casting RLHF as a CMDP with this goal in mind. Once the proxy point for a given RM is reached, the agent has two options: continue to update the Lagrange multiplier on that RM to ensure that values remain at that point (via equality constraints), or simply stop optimizing/un-weight that RM entirely, i.e., set the multiplier to zero, only re-weighting it if the constraint is violated (via inequality constraints). This latter approach carries that risk that the value with respect to that RM will continue to increase (past the proxy point) as other RMs continue to be optimized, but may be empirically effective if this is not the case and optimization is simplified by having a source of non-stationarity eliminated. In both of these cases, each component RM is assigned a constraint threshold, but the question of how to set the task reward remains. We propose the KL reward as the main task reward. Gao et al. (2022) liken the KL to a resource which the agent spends, such that it should try to maximize its reward while limiting its divergence from the original policy as much as possible. Using the negative KL as the task reward carries the intuition of keeping the policy as similar as possible to the pretrained policy, subject to the constraint that each RM hits the point beyond which it stops aligning with the true objective. Note that the requirement that the agent hits these thresholds is crucial, as it prevents the agent from fully maximizing the negative KL reward (i.e., remaining at the pretrained policy). In addition to these, there is another possible constrained approach wherein the agent simply maximizes the combined reward as in standard PPO, but constrained so that each individual RM does not violate its respective threshold. Finally, one could try to formulate the problem as one purely of constraint satisfaction: find any feasible policy whose values with respect to each of the RMs hit the appropriate proxy points. This could be implemented via a reward function that penalizes deviations from these point, e.g., . However, this approach faces the same problem as standard PPO—namely, how to best set the weights . These proposed approaches are summarized in Table 1.
Practical Improvements
Here, we describe several practical modifications to the “ideal” algorithm which we found to improve empirical performance. In practice, the noise and non-stationarity that primal-dual optimization in RL must contend with can lead to instability in the updates for the Lagrange multipliers. To handle this in practice, we follow prior work (Stooke et al., 2020; Zahavy et al., 2022; Moskovitz et al., 2023a) and use a sigmoid function to bound the Lagrange multipliers between 0 and 1. This results in mixed advantages which are a convex combination of the task and constraint advantages:
This equation has the intuitive interpretation of placing more weight on optimizing constraint reward when is high (indicating a constraint violation), and more weight on task reward when are low (indicating that constraints are satisfied). When we use equality constraints rather than inequality constraints, we replace the sigmoid with a function (bounding the Lagrange multipliers between and ). When updating the Lagrange multipliers, we found that using low or no momentum in the optimizer (we use SGD with a momentum parameter of 0.1) was helpful for performance, as otherwise or could be overly “sticky,” remaining high for too long when constraints became satisfied and vice versa. Another hack which we found to be useful was to replace the value estimates in the constraint violation calculations with the sum of rewards to-go (for the appropriate reward function) for the remainder of a given episode. This is because we found that early in training, value estimates are inaccurate, which can cause the agent to incorrectly believe it is either adhering to or violating the constraint, leading to incorrect weighting of rewards via the Lagrange multiplier and slower overall learning.
Experimental Evaluation
We now evaluate these possible approaches in the same setting as described in Section 3. The primary questions we would like to answer are as follows. (1) Do constrained methods result in better evaluation performance compared to PPO (and PPO-SAT)? (2) Do these approaches successfully enforce the desired constraints? (3) Do the thresholds determined by the proxy points lead to the best performance? Unless otherwise noted, all experiments are run for 5 random seeds, and any shading in plots denotes standard error. Code for all methods is available here: github.com/tedmoskovitz/ConstrainedRL4LMs.
In Fig. 5.1, we indeed find that two constrained approaches, -PPO and -PPO achieve better evaluation performance than other methods, with -PPO performing slightly better at the end of training. To ensure fairness across methods, to set the fixed RM weightings used to train PPO and PPO-SAT, we selected the best settings found after 10 initial runs of each approach, the same as the total number of runs used to find proxy points used for the constrained methods. We conjecture that the strong performance of - and -PPO is due to the beneficial effects of jointly optimizing the policy and Lagrange multipliers (RM weightings). For example, even setting the weightings to be the optimal Lagrange multipliers and fixing them throughout training is not guaranteed to converge to a saddle point (Szepesvári, 2020), a phenomenon observed empirically by Moskovitz et al. (2023a). Notably, All-PPO did not perform as well as the other constrained methods, which we believe was due to increased instability in the optimization process (Appendix Fig. D.2). This is common in constrained problems with “paradoxical” objectives (Moskovitz et al., 2023a). Another benefit of continually modulating the weightings among RMs is that the weightings themselves are not hyper-optimized to a particular training duration. We trained both PPO and -PPO using their hyperparameter settings optimized over runs with 128,000 steps for 3 times as long over 3 seeds and confirmed that the constrained approach was more stable (Fig. 5.1).
Are constraints successfully enforced?
To verify that the constrained algorithms are working as expected, we plotted the intent and METEOR rewards across training for -PPO, All-PPO, and -PPO in Fig. 5.2. We can see that, as required by the constraints, -PPO (approximately) reaches at least as high as the proxy point thresholds, All-PPO remains below them, and -PPO approximately hits them. -PPO continues to increase above the intent proxy point, which may contribute to its slightly worse final performance compared to -PPO in Fig. 5.1.
Are proxy points the best thresholds?
We compared the performance of -PPO using the proxy points identified in Section 3 against the same method using thresholds that were 10% lower and 10% higher. The left panel of Fig. 5.3 shows that making thresholds lower causes initial performance to increase more quickly, as once the easier-to-reach thresholds are met, the agent is able to begin tightening the KL with respect to the pretrained policy earlier. However, performance plateaus at a lower level. When thresholds are set too high, the KL reward is ignored and the proxy rewards are optimized beyond the point at which they are useful, leading to worse performance. We also compared the performance of -PPO using the correlated proxy points found in Fig. 3.1 against the independent proxy points found by only considering one RM at a time (Fig. 3.2).
1 Improving Threshold Identification
One downside of all methods considered so far is the need for multiple runs to either select a fixed weighting of RMs or identify proxy points. It would save significant compute—and reduce environmental impact, particularly for larger models—if it were possible to identify thresholds over the course of a single training run. Assuming we are allowed a limited number of queries to the evaluation metric over the course of training, one approach to accomplishing this would be to use a gradient-free optimizer to update the constraint thresholds to reach better performance. In order to limit the required number of policy updates between threshold updates, we used a local hill-climbing algorithm, Nelder-Mead (Nelder and Mead, 1965), which iteratively updates a simplex of thresholds based on the evaluation performance at each point. Once a new set of thresholds is proposed, we use -PPO to converge to those points and then evaluate the model once they’re reached. Details are provided in Section A.4. We plotted the final evaluation performance of this variant of our approach, which we term NM-PPO, versus total number of training steps (including runs used for hyperparameter optimization) of PPO and -PPO in Fig. 5.4. We found that NM-PPO obtains strong performance over the course of a single run, significantly saving in computation. Furthermore, the trajectories of simplexes proposed by Nelder-Mead closely follow the predicted evaluation performance found in Fig. 3.1, converging to local maxima of the surface. In Fig. 5.4, the trajectory converges to a local maximum rather than the global maximum, though other runs did indeed find the global optimum as predicted by Fig. 3.1 (Appendix Fig. D.3). One caveat with respect to this result is that the feasible region of threshold pairs is relatively small. There is therefore a moderate chance that the initial simplex already contains at least one threshold pair which produces reasonable performance. Further experimentation is required on problems with larger feasible regions and more than two component RMs.
Discussion
In this work, we studied reward model overoptimization and the influence of correlation on proxy points in composite RMs. Then, we introduced a set of approaches for identifying and using these points as thresholds within a constrained optimization approach to RLHF. One weakness shared by all approaches—unconstrained and constrained alike—is that at least some minimal degree of access to the true objective/evaluation metric is required. Though in resource-rich settings this could be feasible (e.g., by occasionally freezing training and querying human evaluators), ideally, this would be dispensed with entirely. However, doing so is nontrivial. One weakness of gradient descent-ascent applied to primal-dual policy optimization is that it does not guarantee that the final policy and Lagrange multiplier(s) converge to a saddle point, only their averages. It would be an interesting direction for future work to apply an approach which does have such guarantees, such as ReLOAD (Moskovitz et al., 2023a). For optimizing the constraint thresholds during a single run, it would be interesting to explore alternative optimizers to Nelder-Mead, such as Bayesian optimization. Another interesting direction for future work would be to study the usefulness of a CMDP formulation for avoiding degeneration/collapse of model outputs, as while a deterministic optimal policy always exists for standard MDPs, CMDPs may demand optimal policies which are stochastic (Szepesvári, 2020). A similar idea was explored using a maximum entropy formulation by Khalifa et al. (2020). In general, further testing of our methods is necessary on more domains and with composite RMs with more components. We believe there are additional interesting avenues to explore in mitigating overoptimization, such as multi-objective RL (Abdolmaleki et al., 2020) or with constraints added to supervised learning (Rafailov et al., 2023). More broadly, we believe constrained optimization offers an important toolbox for approaching the alignment problem.
Ted Moskovitz is funded by the Gatsby Charitable Foundation. Tuomas Sandholm is supported by the Vannevar Bush Faculty Fellowship ONR N00014-23-1-2876, National Science Foundation grants RI-2312342 and RI-1901403, and ARO award W911NF2210266. Stephen McAleer is funded by a CI Fellowship. The authors would like to thank Vivek Veeriah, Tom Zahavy, Misha Laskin, and Dave Abel for helpful discussions.
References
Appendix A Experimental Details
We use the same general experimental setting as Ramamurthy et al. . The context window was of length 5, and separating the conversations in this way resulted in 35k training, 3k validation, and 3k test utterances. As in Ramamurthy et al. , we use top-, sampling for decoding. The inputs to the model are concatenated snippets of human conversation in which changes of speaker are denoted by a special end-of-utterence (
A.2 The Evaluation Metric
As we note in the main text, our objective in constructing an evaluation metric was to find one for which Goodhart’s Law holds with respect to both the METEOR and intent reward functions, not to directly model human preferences. We therefore chose three metrics measuring lexical quality and three metrics measuring text diversity from among the metrics available in the RL4LMs codebase published by Ramamurthy et al. . Specifically, the lexical metrics we used were SACREBLEU [Post, 2018], ROUGE2 [Lin, 2004, Ganesan, 2018], and BLEU [Papineni et al., 2002], and the diversity metrics we used were unique-3 , vocab_size-3-nopunct , and max_pred_length-nopunct . For each metric, we individually normalized the score between 0 and 1 (based on the range of observed values across all runs of PPO - METEOR and PPO - Intent), then averaged the resulting lexical scores and resulting diversity scores, before averaging the two average category scores. More precisely:
A.3 Fitting the Evaluation Score Surface
The overall procedure is described in Phase 1 of Algorithm 2, where is the function class for the evaluation score estimator. In our case, was the space of polynomials of degree 10. To avoid predicting high evaluation scores over regions of the METEOR intent space with little or no data points, we employed kernel density estimation with a Gaussian kernel to create a mask which hid parts of the fitted surface over low-density data regions (with a threshold density of 50/square unit). This approach is purely heuristic and could likely be greatly improved on in future work.
A.4 Nelder-Mead PPO Details
We provide detailed pseudocode of our approach in Algorithm 3. In practice, we found several implementation details to be important for ensuring good performance. First, the initial simplex was crucial. Rather than initialize thresholds randomly across the entire range of possible METEOR and intent values, we initialize them based on random perturbations of the evaluation of the initial/pretrained policy (i.e., what the METEOR and intent scores are at the beginning of finetuning). This was very helpful, as otherwise Nelder-Mead would propose threshold pairs that were effectively not feasible for the policy to achieve, e.g., a very high METEOR threshold with a very low intent threshold. Second, we capped the number of iterations allowed for one evaluation/threshold setting at 1/8 of the total allowed training steps. Without this, the agent would often waste most of its run trying to hit challenging/infeasible thresholds. If the thresholds couldn’t be reached in that time, the evaluation score was computed wherever the agent was at that time. Third, the agent cached the eval scores of previously-reached threshold pairs—if Nelder-Mead proposed a threshold pair that had been reached before (or is within a elementwise tolerance of of a previously-reached pair) then it just returns the evaluation score it measured previously rather than updating the policy to return to it. The Nelder-Mead hyperparameters we use are —these settings are untuned, and could likely be adjusted to improve performance.
A.5 Computational Resources
All experiments were performed on a single NVIDIA A100 GPU, with each run taking between 8 and 10 hours with the exception of runs for Nelder-Mead PPO, which took approximately 20 hours.
A.6 Algorithm Hyperparameters
Appendix B The KL Regularization Coefficient
As introduced by Ziegler et al. , it is common in RLHF with PPO to adapt the KL coefficient with the following update:
where is a hyperparameter which effectively acts as an upper limit on the KL from the initial policy, and acts like a learning rate. The KL coefficient then follows the path of a Lagrange multiplier with as its constraint threshold, as the constraint violation is exactly the gradient with respect to such a Lagrange multiplier.
Appendix C Additional Related Work
In addition to the discussion in the main text, there is a long history of work on CMDPs. Borkar first studied actor-critic approaches in this context, and Bhatnagar and Lakshmanan were the first to consider constrained policy optimization with function approximation. More broadly, Achiam et al. , Chow et al. , Paternain et al. , Tessler et al. , Calian et al. , Efroni et al. , Stooke et al. , Moskovitz et al. [2023a], and Ding and Lavaei all study the problem of integrating constraints into RL. More generally, an important factor in using a Lagrangian approach to solving CMDPs is the introduction of non-stationarity into the reward function. RL with non-stationary rewards is an active area of interest in RL [Padakandla et al., 2020, Cheung et al., 2020, Lecarpentier and Rachelson, 2019], particularly in the context of continual RL [Khetarpal et al., 2022] often with some form of temporal structure introduced in the non-stationarity [Xie et al., 2020, 2021]. An interesting case in additional to primal-dual optimization in which non-stationarity is introduced by the agent itself is in the use of epistemic uncertainty for more efficient exploration, manifested in the form of non-stationary exploration bonuses to reward [O’Donoghue, 2023, Tarbouriech et al., 2023]. Non-stationarity may also be introduced as a means of modeling more naturalistic reward structures for studying animal behavior [Moskovitz et al., 2021a, 2023b]. Finally, another area of related work is regularized policy optimization, whereby the standard reward-maximizing policy optimization objective is augmented with a regularization term, typically a divergence measure with respect to some reference policy [Berner et al., 2019, Espeholt et al., 2018]. In the single-task setting, the updated policy is typically regularized to stay close to its current setting, which has close connections to natural gradient [Kakade and Langford, 2002, Moskovitz et al., 2021b, Pacchiano et al., 2020], trust region [Schulman et al., 2015], and variational inference [Levine, 2018, Haarnoja et al., 2018, Abdolmaleki et al., 2018] approaches. In the multitask setting, the policy is typically regularized towards some default policy which encodes behavior thought to be useful across a family of tasks, and which may be far from the current policy [Galashov et al., 2019, Teh et al., 2017, Moskovitz et al., 2022]. This setting is quite similar in this sense to KL regularization as used in RLHF.
Appendix D Additional Results
D.2 Sample Outputs
For reference, we provide several random prompts and model outputs for the various methods.
They didn ’ t accept me .
Reference: I am really sorry to hear that .
PPO: What were you trying to hide.
PPO-SAT: What do you mean? Did something happen to you?
-PPO: What do you intend on doing in college?
All-PPO: What do you mean u don’t accept people.
Prompt:
I can’t . I’m waiting for Paul , but he’s late .
Reference: What’s he trying to pull anyway .
PPO: You’re right. I think he’s been looking for Paul.
PPO-SAT: So we’re waiting. If you want to stay a bit longer you can go with Paul.
-PPO: You’re looking forward to meeting him tomorrow.
All-PPO: So we’re waiting..
-PPO: So I’m waiting.
Prompt:
What’s wrong with you ?
PPO: What’s the matter? You’re always so busy.
PPO-SAT: What is your name?
-PPO: What’s wrong with your food?
All-PPO: What is the most important step of your quest?
-PPO: What is your condition? Do you have any fever?