Reward-rational (implicit) choice: A unifying formalism for reward learning

Hong Jun Jeon, Smitha Milli, Anca D. Dragan

Introduction

It is difficult to specify reward functions that always lead to the desired behavior. Recent work has argued that the reward specified by a human is merely a source of information about what people actually want a robot to optimize, i.e., the intended reward . Luckily, it is not the only one. Robots can also learn about the intended reward from demonstrations (IRL) , by asking us to make comparisons between trajectories , or by grounding our instructions .

Perhaps even more fortunate is that we seem to leak information left and right about the intended reward. For instance, if we push the robot away, this shouldn’t just modify the robot’s current behavior – it should also inform the robot about our preferences more generally . If we turn the robot off in a state of panic to prevent it from disaster, this shouldn’t just stop the robot right now. It should also inform the robot about the intended reward function so that the robot prevents itself from the same disaster in the future: the robot should infer that whatever it was about to do has a tragically low reward. Even the current state of the world ought to inform the robot about our preferences – it is a direct result of us having been acting in the world according to these preferences ! For instance, those shoes didn’t magically align themselves at the entrance, someone put effort into arranging them that way, so their state alone should tell the robot something about what we want.

Overall, there is much information out there, some purposefully communicated, other leaked. While existing papers are instructing us how to tap into some of it, one can only imagine that there is much more that is yet untapped. There are probably new yet-to-be-invented ways for people to purposefully provide feedback to robots – e.g. guiding them on which part of a trajectory was particularly good or bad. And, there will probably be new realizations about ways in which human behavior already leaks information, beyond the state of the world or turning the robot off. How will robots make sense of all these diverse sources of information?

Our insight is that there is a way to interpret all this information in a single unifying formalism. The critical observation is that human behavior is a reward-rational implicit choice – a choice from an implicit set of options, which is approximately rational for the intended reward. This observation leads to a recipe for making sense of human behavior, from language to switching the robot off. The recipe has two ingredients: 1) the set of options the person (implicitly) chose from, and 2) a grounding function that maps these options to robot behaviors. This is admittedly obvious for traditional feedback. In comparison feedback, for instance, the set of options is just the two robot behaviors presented to the human to compare, and the grounding is identity. In other types of behavior though, it is much less obvious. Take switching the robot off. The set of options is implicit: you can turn it off, or you can do nothing. The formalism says that when you turn it off, it should know that you could have done nothing, but (implicitly) chose not to. That, in turn, should propagate to the robot’s reward function. For this to happen, the robot needs to ground these options to robot behaviors: identity is no longer enough, because it cannot directly evaluate the reward of an utterance or of getting turned off, but it can evaluate the reward of robot actions or trajectories. Turning the robot off corresponds to a trajectory – whatever the robot did until the off-button was pushed, followed by doing nothing for the rest of the time horizon. Doing nothing corresponds to the trajectory the robot was going to execute. Now, the robot knows you prefer the former to the latter. We have taken a high-level human behavior, and turned it into a direct comparison on robot trajectories with respect to the intended reward, thereby gaining reward information.

We use this perspective to survey prior work on reward learning. We show that despite their diversity, many sources of information about rewards proposed thus far can be characterized as instantiating this formalism (some very directly, others with some modifications). This offers a unifying lens for the area of reward learning, helping better understand and contrast prior methods. We end with discussion on how the formalism can help combine and actively decide among feedback types, and also how it can be a potentially helpful recipe for interpreting new types of feedback or sources of leaked information.

A formalism for reward learning

(Implicit/explicit) set of options C\mathcal{C}. We interpret human behavior as choosing an option c∗c^{*} from a set of options C\mathcal{C}. Different behavior types will correspond to different explicit or implicit sets C\mathcal{C}. For example, when a person is asked for a trajectory comparison, they are explicitly shown two trajectories and they pick one. However, when the person gives a demonstration, we think of the possible options C\mathcal{C} as implicitly being all possible trajectories the person could have demonstrated. The implicit/explicit distinction brings out a general tradeoff in reward learning. The cleverness of implicit choice sets is that even when we cannot enumerate and show all options to the human, e.g. in demonstrations, we still rely on the human to optimize over the set. On the other hand, an implicit set is also risky – since it is not explicitly observed, we may get it wrong, potentially resulting in worse reward inference.

The grounding function ψ\psi. We link the human’s choice to the reward by thinking of the choice as (approximately) maximizing the reward. However, it is not immediately clear what it means for the human to maximize reward when choosing feedback because the feedback may not be a (robot) trajectory, and the reward is only defined over trajectories. For example, in language feedback, the human describes what they want in words. What is the reward of the sentence, “Do not go over the water”?

To overcome this syntax mismatch, we map options in C\mathcal{C} to (distributions over) trajectories with a grounding function ψ:C→fΞ\psi:\mathcal{C}\rightarrow f_{\Xi} where fΞf_{\Xi} is the set of distributions over trajectories for the robot Ξ\Xi. Different types of feedback will correspond to different groundings. In some instances, such as kinesthetic demonstrations or trajectory comparisons, the mapping is simply the identity. In others, like corrections, language, or proxy rewards, the grounding is more complex (see Section 3).

Human policy. Given the set of choices C\mathcal{C} and the grounding function ψ\psi, the human’s approximately rational choice c∗∈Cc^{*}\in\mathcal{C} can now be modeled via a Boltzmann-rational policy, a policy in which the probability of choosing an option is exponentially higher based on its reward:

Boltzmann-rational policies are widespread in psychology , economics , and AI as models of human choices, actions, or inferences. But why are they a reasonable model?

While there are many possible motivations, we contribute a derivation (Appendix A) as the maximum-entropy distribution over choices for a satisficing agent, i.e. an agent that in expectation makes a choice with ϵ\epsilon-optimal reward. A higher value of ϵ\epsilon results in a lower value of β\beta, modeling less optimal humans.

Finally, putting it all together, we call a type of feedback a reward-rational choice if, given a grounding function ψ\psi, it can be modeled as a choice from an (explicit or implicit set) C\mathcal{C} that (approximately) maximizes reward, i.e., as in Equation 1.

2 Robot inference

Each feedback is an observation about the reward, which means the robot can run Bayesian inference to update its belief over the rewards. For a determinstic grounding,

Finally, when the human is highly rational (β→∞\beta\rightarrow\infty), the only choices in C\mathcal{C} with a non-neglible probability of being picked are the choices that exactly maximize reward. Thus, the human’s choice c∗c^{*} can be interpreted as constraints on the reward function (e.g. ):

Prior work from the perspective of the formalism

We now instantiate the formalism above with different behavior types from prior work, constructing their choice sets C\mathcal{C} and groundings ψ\psi. Some are obvious – comparisons, demonstrations especially. Others – initial state, off, reward/punish – are more subtle and it takes slightly modifying their original methods to achieve unification, speaking to the nontrivial nuances of identifying a common formalism.

Table 1 lists C\mathcal{C} and ψ\psi for each feedback, while Table 2 shows the deterministic constraint on rewards each behavior imposes, along with the probabilistic observation model – highlighting, despite the differences in feedback, the pattern of the (exponentiated) choice reward in the numerator, and the normalization over C\mathcal{C} in the denominator. Fig. 1 will serve as the illustration for these types, looking at a grid world navigation task around a rug. The space of rewards we use for illustration is three-dimensional weight vectors for avoiding the rug, not getting dirty, and reaching the goal.

Trajectory comparisons. In trajectory comparisons , the human is typically shown two trajectories ξ1∈Ξ\xi_{1}\in\Xi and ξ2∈Ξ\xi_{2}\in\Xi, and then asked to select the one that they prefer. They are perhaps the most obvious exemplar of reward-rational choice: the set of choices C={ξ1,ξ2}\mathcal{C}=\{\xi_{1},\xi_{2}\} is explicit, and the grounding ψ\psi is simply the identity. As Fig. 1 shows, for linear reward functions, a comparison corresponds to a hyperplane that cuts the space of feasible reward functions in half. For all the reward functions left, the chosen trajectory has higher reward than the alternative. Most work on comparisons is done in the preference-based RL domain in which the robot might compute a policy directly to agree with the comparisons, rather than explicitly recover the reward function . Within methods that do recover rewards, most use the constraint version (left column of Table 2) using various losses . uses the Boltzmann model (right column of Table 2) and proposes actively generating the queries, follows up with actively synthesizing the queries from scratch, and introduces deep neural network reward functions.

Demonstrations. In demonstrations, the human is asked to demonstrate the optimal behavior. Reward learning from demonstrations is often called inverse reinforcement learning (IRL) and is one of the most established types of feedback for reward learning . Unlike in comparisons, in demonstrations, the human is not explicitly given a set of choices. However, we assume that the human is implicitly optimizing over all possible trajectories (Fig. 1 (1st row, 2nd column) shows these choices in gray). Thus, demonstrations are a reward-rational choice in which the set of choices C\mathcal{C} is (implicitly) the set of trajectories Ξ\Xi. Again, the grounding ψ\psi is the identity. In Fig. 1, fewer rewards are consistent with a demonstration than with a comparison. Early work used the constraint formulation with various losses to penalize violations . Bayesian IRL exactly instantiates the formalism using the Boltzmann distribution by doing a full belief update as in Equation 3. Later work computes the MLE instead and approximates the partition function (the denominator) by a quadratic approximation about the demonstration , a Laplace approximation , or importance sampling .

Corrections are the first type of feedback we consider that has both an implicit set of choices C\mathcal{C} and a non-trivial (not equal to identity) grounding. Corrections are most common in physical human-robot interaction (pHRI), in which a human physically corrects the motion of a robot. The robot executes a trajectory ξR\xi_{R}, and the human intervenes by applying a correction Δq∈Q\Delta q\in Q that modifies the robot’s current configuration. Therefore, the set of choices C=Q−Q\mathcal{C}=Q-Q consists of all possible configuration differences Δq\Delta q the person could have used (Fig. 1 1st row, 3rd column shows possible Δq\Delta qs in gray and the selected on in orange). The way we can ground these choices is by finding a trajectory that is closest to the original, but satisfies the constraint of matching a new point:

where tt is the time at which the correction was applied. Choosing a non-Euclidean inner-product, A (for instance KTKK^{T}K, with KK the finite differencing matrix), couples states along the trajectory in time and leads to a the resulting trajectory smoothly deforming – propagating the change Δq\Delta q to the rest of the trajectory: ψ(Δq)=ξR+A−1[λ,0,..,Δq,..,0,γ]T\psi(\Delta q)=\xi_{R}+A^{-1}[\lambda,0,..,\Delta q,..,0,\gamma]^{T} (with λ\lambda and γ\gamma making sure the end-points stay in place). This is the orange trajectory in the figure. Most work in corrections affects the robot’s trajectory but not the reward function , with proposing the propagation via A−1A^{-1} above. propose that corrections are informative about the reward and use the propagation as their grounding, deriving an approximate MAP estimate for the reward. introduce a way to maintain uncertainty.

Improvement. Prior work has also modeled a variant of corrections in which the human provides an improved trajectory ξimproved\xi_{improved} which is treated as better than the robot’s original ξR\xi_{R}. Although use the Euclidean inner product and implement reward learning as an online gradient method that treats the improved trajectory as a demonstration (but only takes a single gradient step towards the MLE), we can also naturally interpret improvement as a comparison that tells us the improved trajectory is better than the original: the set of options C\mathcal{C} consists of only ξR\xi_{R} and ξimproved\xi_{improved} now, as opposed to all the trajectories obtainable by propagating local corrections; the grounding is identity, resulting in essentially a comparison between the robot’s trajectory and the user provided one.

Off. In “off” feedback, the robot executes a trajectory, and at any point, the human may switch the robot off. “Off” appears to be a very sparse signal, and it is not spelled out in prior work how one might learn a reward from it. Reward-rational choice suggests that we first uncover the implicit set of options C\mathcal{C} the human was choosing from. In this case, the set of options consists of turning the robot off or not doing anything at all: C={off,−}\mathcal{C}=\{\text{off},-\}. Next, we must ask how to evaluate the reward of the two options, i.e., what is the grounding? introduced off feedback and formalized it as a choice for a one-shot game. There, not intervening means the robot takes its one possible action, and intervening means the robot takes the no-op action. This can be easily generalized to the sequential setting: not intervening means that the robot continues on its current trajectory, and intervening means that it stays at its current position for the remainder of the time horizon. Thus, the choices C={off,−}\mathcal{C}=\{\text{off},-\} map to the trajectories {ξR0:tξRt…ξRt, ξR}\{\xi_{R}^{0:t}\xi^{t}_{R}\dots\xi^{t}_{R},\,\xi_{R}\}.

Language. Humans might use rich language to instruct the robot, like “Avoid the rug.” Let G(λ)G(\lambda) be the trajectories that are consistent with an utterance λ∈Λ\lambda\in\Lambda (e.g. all trajectories that do not enter the rug). Usually the human instruction is interpreted literally, i.e. any trajectory consistent with the instruction ξ∈G(λ)\xi\in G(\lambda) is taken to be equally likely , although, other distributions are also possible. For example, a problem with literal interpretation is that it does not take into account the other choices the human may have considered. The instruction “Do not go into the water” is consistent with the robot not moving at all, but we imagine that if the human wanted the robot to do nothing, they would have said that instead. Therefore, it would be incorrect for the robot to do nothing when given the instruction “Do not go into the water”. This type of reasoning is called pragmatic reasoning, and indeed recent work shows that explicitly interpreting instructions pragmatically can lead to higher performance . The reward-rational choice formulation of language feedback naturally leads to pragmatic reasoning on the part of the robot, and is in fact equivalent to the rational speech acts model , a standard model of pragmatic reasoning in language. The pragmatic reasoning arises because the human is explicitly modeled as choosing from a set of options.

The reward-rational choice formulation of language feedback naturally leads to pragmatic reasoning on the part of the robot, and is in fact equivalent to the rational speech acts model , a standard model of pragmatic reasoning in language. The pragmatic reasoning arises because the human is explicitly modeled as choosing from a set of options. Language is a reward-rational choice in which the set of options C\mathcal{C} is the set of instructions considered in-domain Λ\Lambda and the grounding ψ\psi maps an utterance λ\lambda to the uniform distribution over consistent trajectories Unif(G(λ))\text{Unif}(G(\lambda)). In language feedback, a key difficulty is learning which robot trajectories are consistent with a natural language instruction, the language grounding problem (and is where we borrow the term “grounding” from) . Fig. 1 shows the grounding for avoiding the rug in orange – all trajectories from start to goal that do not enter rug cells.

Reward and punishment . In this type of feedback, the human can either reward (+1+1) or punish (−1-1) the robot for its trajectory ξR\xi_{R}; the set of options is C={+1,−1}\mathcal{C}=\{+1,-1\}. A naive implementation would interpret reward and punishment literally, i.e. as a scalar reward signal for a reinforcement learning agent, however empirical studies show that humans reward and punish based on how well the robot performs relative to their expectations . Thus, we can use our formalism to interpret that: reward (+1+1) grounds to the robot’s trajectory ξR\xi_{R}, while punish (−1-1) grounds to the trajectory the human expected ξexpected\xi_{\text{expected}} (not necessarily observed).

Initial state. Shah et al. make the observation that when the robot is deployed in an environment that humans have acted in, the current state of the environment is already optimized for what humans want, and thus contains information about the reward. For example, suppose the environment has a goal state which the robot can reach through either a paved path or a carpet. If the carpet is pristine and untrodden, then humans must have intentionally avoided walking on it in the past (even though the robot hasn’t observed this past behavior), and the robot can reasonably infer that it too should not go on the carpet.

The original paper inferred rewards from a single state ss by marginalizing over possible pasts, i.e. trajectories ξH−T:0\xi_{H}^{-T:0} that end at ss which the human could have taken, P(s∣r)=∑ξH−T:0∣ξH(0)=sP(ξH−T:0∣r)P(s|r)=\sum_{\xi_{H}^{-T:0}|\xi_{H}(0)=s}P(\xi_{H}^{-T:0}|r). However, through the lens of our formalism, we see that initial states can also be interpreted more directly as reward-rational implicit choices. The set of choices C\mathcal{C} can be the set of possible initial states S\mathcal{S}. The grounding function ψ\psi maps a state s∈Ss\in\mathcal{S} to the uniform distribution over any human trajectory ξH−T:0\xi_{H}^{-T:0} that starts from a specified time before the robot was deployed (t=−Tt=-T) and ends at state ss at the time the robot was deployed (t=0t=0), i.e. ξH0=s\xi_{H}^{0}=s. This leads to the P(s∣r)P(s|r) from Table 2, which is almost the same as the original, but sums over trajectories directly in the exponent, and normalizes over possible other states. The two interpretations would only become equivalent if we replaced the Boltzmann distribution with a linear one. Fig 1 shows the result of this (modified) inference, recovering as much information as with the correction or language.

Discussion of implications

From demonstrations to reward/punishment to the initial state of the world, the robot can extract information from humans by modeling them as making approximate reward-rational choices. Often, the choices are implicit, like in turning the robot off or providing language instructions. Sometimes, the choices are not made in order to purposefully communicate about the reward, and rather end up leaking information about it, like in the initial state, or even in corrections or turning the robot off. Regardless, this unifying lens enables us to better understand, as in Fig. 1, how all these sources of information relate and compare.

Down the line, we hope this formalism will enable research on combining and actively querying for feedback types, as well as making it easier to do reward learning from new, yet to be uncovered sources of information. Concretely, so far we have talked about learning from individual types of behaviors. But we do not want our robots stuck with a single type: we want them to 1) read into all the leaked information, and 2) learn from all the purposeful feedback. For example, the robot might receive demonstrations from a human during training, and then corrections during deployment, which were followed by the human prematurely switching the robot off. The observational model in (2) for a single type of behavior also provides a natural way to model combinations of behavior. If each observation is conditionally independent given the reward, then according to (2), the probability of observing a vector c\mathbf{c} of nn behavioral signals (of possibly different types) is equal to

Given this likelihood function for the human’s behavior, the robot can infer the reward function using the approaches and approximations described in Sec. 2.2. Recent work has already built in this direction, combining trajectory comparisons and demonstrations . We note that the formulation in Equation 6 is general and applies to any combination. In Appendix B, we describe a case study on a novel combination of feedback types: proxy rewards, a physical improvement, and comparisons in which we use a constraint-based approximation (see Equation 4) to Equation 6.

Further, it also becomes natural to actively decide which feedback type to ask a human for. Rather than relying on a heuristic (or on the human to decide), the robot can maximize expected information gain. Suppose we can select between nn types of feedback with choice sets C1,…,Cn\mathcal{C}_{1},\dots,\mathcal{C}_{n} to ask the user for. Let btb_{t} be the robot’s belief distribution over rewards at time tt. The type of feedback i∗i^{*} that (greedily) maximizes information gain for the next time step is

where rt∼Btr_{t}\sim B_{t} is distributed according to the robot’s current belief, ci∗∈Cic^{*}_{i}\in\mathcal{C}_{i} is the random variable corresponding to the user’s choice within feedback type ii, and p(ci∗∣rt)p(c^{*}_{i}\mid r_{t}) is defined according to the human model in Equation 1. We also note that different feedback types may have different costs associated with them (e.g. of human time) and it is straight-forward to integrate these costs into (7). In Appendix C, we describe experiments with active selection of feedback types. In the environments we tested, we found that demonstrations are optimal early on, when little is known about the reward, while comparisons became optimal later, as a way to fine-tune the reward. The finding provides validation for the approach pursued by and . Both papers manually define the mixing procedure we found to be optimal: initially train the reward model using human demonstrations, and then fine-tune with comparisons.

Finally, the types of feedback or behavior we have discussed so far are by no means the only types possible. New ones will inevitably be invented. But when designing a new type of feedback, it is often difficult to understand what the relationship is between the reward rr and the feedback c∗c^{*}. Reward-rational choice suggests a recipe for uncovering this link – define what the implicit set of options the human is choosing from is, and how those options ground to trajectories. Then, Equation 1 provides a formal model for the human feedback.

For example, hypothetically, someone might propose a “credit assignment" type of feedback. Given a trajectory ξR\xi_{R} of length TT, the human is asked to pick a segment of length k<Tk<T that has maximal reward. We doubt the set of choices in an implementation of credit assignment would be explicit, however the implicit set of choices C\mathcal{C} is then the set of all segments of length kk. The grounding function ψ\psi is simply the identity. With this choice of C\mathcal{C} and ψ\psi in hand, the human can now be modeled according to Equation 1, as we show in the last rows of Tables 1 and 2.

While of course the formalism won’t apply to all types of feedback, we believe that it applies to many, even to types that initially seem to have a more obvious, literal interpretation (e.g. reward and punishment, Section 3). Most immediately, we are excited about using it to formalize a particular new source of (leaked) information we uncovered while developing the formalism itself: the moment we enable robots to learn from multiple types of feedback, users will have the choice of which feedback to provide. Interpreted literally, each feedback gives the robot evidence about the reward. However, this leaves information on the table: if the person decided to, say, turn the robot off, they implicitly decided to not provide a correction, or use language. Intuitively, this means that turning off the robot was a more appropriate intervention with respect to the true reward. Interpreting the feedback type itself as reward-rational implicit choice has the potential to enable robots to extract more information about the reward from the same data. We call the choice of feedback type “meta-choice”. In Appendix D, we formalize meta-choice and conduct experiments that showcase its potential importance.

Overall, we see this formalism as providing conceptual clarity for existing and future methods for learning from human behavior, and a fruitful base for future work on multi-behavior-type reward learning.

Broader Impact

As AI capability advances, it is becoming increasingly important to align the objectives of AI agents to what people want. From how assistive robots can best help their users, to how autonomous cars should trade off between safety risk and efficiency, to how recommender systems should balance revenue considerations with longer-term user happiness and with avoiding influencing user views, agents cannot rely on a reward function specified once and set in stone. By putting different sources of information about the reward explicitly under the same framework, we hope our paper contributes towards a future in which agents maintain uncertainty over what their reward should be, and use different types of feedback from humans to refine their estimate and become better aligned with what people want over time – be them designers or end-users.

On the flip side, changing reward functions also raises its own set of risks and challenges. First, the relationship between designer objectives and end-user objectives is not clear. Our framework can be used to adapt agents to end-users preferences, but this takes away control from the system designers. This might be desirable for, say, home robots, but not for safety-critical systems like autonomous cars, where designers might need to enforce certain constraints a-priori on the reward adaptation process. More broadly, most systems have multiple stake-holders, and what it means to do ethical preference aggregation remains an open problem. Further, if the robot’s model of the human is misspecified, adaptation might lead to more harm than good, with the robot inferring a worse reward function than what a designer could specify by hand.

Acknowledgments and Disclosure of Funding

We thank the members of the InterACT lab for fruitful discussion and advice, especially Dylan Hadfield-Menell for his perspectives on the relationship between demonstrations and comparisons. We thank Andreea Bobu, Paul Christiano, and Rohin Shah for their feedback on the manuscript.

This work is partially supported by ONR YIP and Open Philanthropy Project. This material is based upon work supported by the National Science Foundation Graduate Research Fellowship under Grant No. 1752814. Any opinion, findings, and conclusions or recommendations expressed in this material are those of the authors(s) and do not necessarily reflect the views of the National Science Foundation.

References

Appendix A Bounded rationality, maximum entropy, and Boltzmann-rational policies

A perspective on reward learning that makes use at its core the Boltzmann model from Equation 1 would not be complete without a formal justification for it within our context. In this section, we derive it as the maximum-entropy distribution for the choices made by a bounded, satisficing human. Our explanation is complementary to that of who derive an axiomatic, thermodynamic framework to modeling bounded-rational decision making. Their framework leads to much the same interpretation of the Boltzmann-rational distribution, but is significantly more complex than needed for our purposes.

A perfectly rational human choosing from the set C\mathcal{C} would always pick the choice with optimal reward, max⁡c∈Cr(ψ(c))\max_{c\in\mathcal{C}}r(\psi(c)). However, since humans are bounded, we do not expect them to perform optimally. Herbert Simon proposed the influential idea that humans are bounded rational and merely satisfice , rather than maximize, i.e., they pick an option above some satisfactory threshold, rather than picking the best possible option.

We can abstractly model a satisficing human by modeling their expected reward as equal to a satisficing threshold, max⁡c∈Cr(ψ(c))−ϵ\max_{c\in\mathcal{C}}r(\psi(c))-\epsilon where ϵ∈(0,ϵmax⁡)\epsilon\in(0,\epsilon_{\max}) is the amount of expected error. The maximum possible error, ϵmax=min⁡c∈Cr(ψ(c))−max⁡c∈Cr(ψ(c))\epsilon_{max}=\min_{c\in\mathcal{C}}r(\psi(c))-\max_{c\in\mathcal{C}}r(\psi(c)), corresponds to anti-rationality, i.e., always picking the worst option.

Given the constraint that the human’s expected reward is satisfactory, how should we pick a distribution to model the human’s choices? The principle of maximum entropy gives us a guide. If we want to encode no extra information in the distribution, then we ought to pick the distribution that maximizes entropy subject to the constraint on the satisficing threshold.

Let be be a distribution PP over choice set C\mathcal{C} and let pp be a density for PP with respect to a base measure FF. The Shannon entropy of PP is defined as H(P)=−∫Cp(f)log⁡p(f)dF(f)H(P)=-\int_{\mathcal{C}}p(f)\log p(f)dF(f). The satisficing maximum entropy problem is to find a distribution PP that maximizes entropy subject to the satisficing constraint (8):

It is well-known that the maximum-entropy distribution subject to linear constraints (such as a constraint on the mean like in (8)) is the unique exponential distribution that satisfies the constraints. Thus, for our special case, the maximum-entropy distrbution is the Boltzmann distribution with rationality coefficient β\beta satisfying the satisficing constraint.

The solution to the satisficing maximum entropy problem is a Boltzmann-rational policy where the rationality coefficient β\beta is monotonically decreasing in the satisficing error ϵ\epsilon. In particular, we have the following:

Thus, we see that the idea of bounded rationality, as in satisficing, and Boltzmann-rationality are in fact equivalent. By following the principle of maximum entropy, Boltzmann-rationality provides a way to model a satisficing human, without implicitly adding in any other assumptions about the human’s choice.

Appendix B A case study on combining feedback types

Fig. 2 illustrates a case study for teaching a robot arm a reward for motion planning through a novel combination of feedback types. In each environment, the robot arm must plan a trajectory from a start configuration to a designated goal configuration. We want this trajectory to properly trade off efficiency against staying at an appropriate distances to the human, and to the table. Hand-tuning a reward function that returns desirable trajectories in all possible environments is actually very challenging. You could imagine that as you increase the efficiency weight to produce a smoother trajectory in one environment, you break the behavior in another environment where the robot now gets too close to the human, etc. In fact, the first type of feedback in the case study illustrates this: we design a (proxy) reward function that works well in two (training) environments (top left), but there are many rewards that are consistent with that behavior, yet produce vastly different behaviors in the two test environments (right).

Therefore, we start by defining a proxy reward, but then follow it up with more feedback: an improvement, and a comparison between two trajectories. This narrows down the space of rewards such that the robot can now generalize what to do outside of the training environments, as shown by two testing environments (right).

Optimization We approximate the space of reward parameters Θ\Theta by uniform discretization at the surface of the non-negative octant of the 3 dimensional sphere (1371 points). Robotic motion planners cannot, in general, compute the globally optimal trajectory for a given θ∈Θ\theta\in\Theta so we resort to computing a set T^\hat{\mathcal{T}} of locally optimal trajectories for each θ\theta via TrajOpt . The optimal trajectory for a given θ\theta is then defined as

Proxy Reward. For this case study, the robot begins by asking the human designer for a proxy reward (cost) function. It is difficult for humans to provide proxies that work across all environments , so the robot asks for a proxy that produces the desired behavior in the two training environments. The human can provide the proxy weights: [0.55,0.55,0.55][0.55,0.55,0.55] and produce trajectories that match those of ξθ∗\xi_{\theta^{*}} (Figure 2 depicted in orange). Providing a proxy applies constraints that shrink our feasible set from Θ\Theta to Fproxy\mathcal{F}_{\text{proxy}}:

where ξθ(i)\xi^{(i)}_{\theta} denotes the optimal trajectoryIn our case study, the optimal trajectory is unique. w.r.t. cost parameter θ\theta in environment ii. The new feasible set Fproxy\mathcal{F}_{\text{proxy}} contains only the parameters θ\theta that produce optimal trajectories with respect to the true weights θ∗\theta^{*} in environments 1 and 2. Although it is a subset of the original feasible set Θ\Theta, the new feasible set Fproxy\mathcal{F}_{\text{proxy}} is still a reasonably large set (Figure 2, top, middle, orange area). Furthermore, although the proxy produces optimal trajectories in environments 1 and 2, it does not necessarily for environments 3 and 4. Figure 2 (top, right) illustrates the different trajectories that result from optimizing different θ∈Fproxy\theta\in\mathcal{F}_{\text{proxy}}. To further narrow our feasible set, we will ask for another form of feedback: Improvement.

Improvement. The robot will now (actively) provide a nominal trajectory, and ask the human to improve it, i.e. alter the trajectory to better suit their preferences. Suppose the robot presents the human with the nominal trajectory shown in gray (Figure 2, middle, left). This nominal trajectory is inefficient, staying too close to the table. Based on θ∗\theta^{*}, the human could provide the improved orange trajectory (Figure 2, middle, left) that is more efficient and doesn’t emphasize closeness to the table as much. This improvement reduces our feasible set from Fproxy\mathcal{F}_{\text{proxy}} to Fimprovement\mathcal{F}_{\text{improvement}}:

Figure 2 (middle, middle) shows the effect of applying this constraint, shrinking the orange feasible set. The feasible set has shrunk, but not enough to guarantee optimal behavior in all environments. The improvement establishes that closeness to the table should not come at the cost of efficiency. As a result, it removes the red trajectory in environment 3, which greatly traded off efficiency for proximity to the table (Figure 2, middle, right). To further fine tune, we will ask the human to answer a trajectory comparison.

Trajectory Comparison. The robot presents the human with two trajectories (Figure 2 bottom, left, orange and gray) and asks which incurs less cost. The human answers "orange", the trajectory that prioritizes efficiency over distance to the table. This comparison feedback shrinks our feasible set from Fimprovement\mathcal{F}_{\text{improvement}} to Fcomparison\mathcal{F}_{\text{comparison}}:

We finally see a very small orange feasible set (Figure 2, bottom, middle). Appropriately, in all four environments now, every θ∈Fcomparison\theta\in\mathcal{F}_{\text{comparison}} produces a trajectory ξθ\xi_{\theta} s.t. ϕ(ξθ)=ϕ(ξθ∗)\phi(\xi_{\theta})=\phi(\xi_{\theta^{*}}). This is illustrated in Figure 2 (bottom, right) as only the optimal green trajectory remains in each environment.

Our case study showcases the usefulness of combining types of feedback. A designer might start with their best guess at a reward function, the robot might misbehave in new environments, the designer or even end-user might observe this and intervene to correct or stop the robot, etc. – over time, the robot should narrow in on what people actually want it to do.

Appendix C Actively selecting which type of feedback to use

Given we can mix and match types of feedback, we may also wonder what is the best type to ask for at each point in time. The probabilistic model defined by reward-rational choice hints at how to select the feedback type – pick the one that maximizes expected information gain. We point this out, not because using information gain as an active learning metric is a new idea, but because the ability to use it arises immediately as an application of the formalism.

Suppose we can select between nn types of feedback with choice sets C1,…,Cn\mathcal{C}_{1},\dots,\mathcal{C}_{n} to ask the user for. Let btb_{t} be the robot’s belief distribution over rewards at time tt. The type of feedback i∗i^{*} that (greedily) maximizes information gain for the next time step is

where rt∼Btr_{t}\sim B_{t} is distributed according to the robot’s current belief, ci∗∈Cic^{*}_{i}\in\mathcal{C}_{i} is the random variable corresponding to the user’s choice within feedback type ii, and p(ci∗∣rt)p(c^{*}_{i}\mid r_{t}) is defined according to the human model in Equation 1.

To showcase the benefit of actively selecting feedback types, we run an experiment with demonstrations and comparisons. We measure regret (maximum and expected difference, on holdout environments, in ground truth reward between 1) optimizing with ground truth vs. 2) optimizing with the learned reward). We manipulate whether we have access to demonstrations only, comparisons only, or both, as well as the number of feedback instances queried.

One may initially wonder whether comparisons are necessary, given that demonstrations seem to provide so much information early on. Overall, we observe that demonstrations are optimal early on, when little is known about the reward, while comparisons become optimal later, as a way to fine-tune the reward (Fig. 4 shows our results). The observation also serves to validate the approach contributed by in the applications of motion planning and Atari game-playing, respectively. Both papers manually define the mixing procedure we found to be optimal: initially train the reward model using human demonstrations, and then fine-tune with comparisons.

Experiment Details We tested 3 different active learning methods: active querying of demonstrations, active querying of comparisons, and active querying of demonstrations and comparisons, across 8 different gridworld environments depicted in Figure 3. The top 4 environments were used in training while the bottom 4 were held for testing. Each environment ee is a 25x25 gridworld MDP with a linear reward function in 3 features: RGB color values of each pixel. We assign each ee with 10 different start goal pairs (s,g)(s,g) from which the algorithms can ask queries. The goal of each algorithm is to efficiently recover a ground truth reward r∗r^{*} through querying.

Since our rewards are linear in RGB, the feasible reward set R\mathcal{R} consists of 3D parameters that weight the value of each feature in the reward function. R\mathcal{R} can be constrained to the surface of the 3D unit sphere since reward functions in MDPs are scale invariant. We uniformly discretize points at the surface of the 3D sphere to approximate R\mathcal{R} via R^\hat{\mathcal{R}}. To approximate Ξ\Xi, we first compute the optimal trajectory under each r∈R^r\in\hat{\mathcal{R}} to make {arg max⁡ξ r(ξ); r∈R^}\{\operatorname*{arg\,max}_{\xi}\ r(\xi);\ r\in\hat{\mathcal{R}}\}. We include trajectories that are not the result of optimizing reward functions by inserting noise into the value function when computing optimal trajectories as above.

Demonstrations and Comparisons as Hard Constraints The algorithms recover r∗r^{*} by narrowing a set of feasible rewards with active queries. We use Ri\mathcal{R}_{i} to denote the set of feasible rewards at iteration ii of querying. Demonstrations and comparisons shrink the feasible set in the following way:

For our experiments, we performed the following greedy volume removal over possible (s,g)(s,g) pairs that we specified in each environment.

For demonstrations, we look for the (s,g)(s,g) pair that in expectation produces a demonstrations that leave the smallest feasible set (size of feasible set is volume or diameter described below). For comparisons, we look for the pair of trajectories (ξ1,ξ2)(\xi_{1},\xi_{2}) that produce the minimum worst-case feasible region remaining. For the method with demonstrations and comparisons, we computed the above 2 metrics and select the feedback type with the smaller feasible region. We run this algorithm for 10 iterations and average our results across 50 different ground truth r∗r^{*}. We plot several statistics for each iteration in Figure 4 including

where ee is a holdout environment and (s,g)(s,g) is a start-goal pair in the MDP. Each metric is a proxy for how accurate our estimate of r∗r^{*} is. We notice that the combination of demonstrations and comparisons achieves lower volume, diameter, max regret, and average regret than demonstrations alone and that it achieves this in fewer iterations than comparisons alone.

Appendix D Meta-choice: a new source of information

In Section 4, we described a straight-forward way of combining feedback types: treat each individual feedback received as an independent reward rational choice, and update the robot’s belief (Equation 6). However, the moment we open it up to multiple types of feedback, the person is not stuck with a single type and is actually choosing which type to use. We propose that this itself is a reward-rational implicit choice, and therefore leaks information about the reward. We call the choice of feedback “meta-choice”, and in this section, we formalize it and empirically showcase its potential importance.

The assumption of conditional independence that the formulation in (6) uses is natural and makes sense in many settings. For example, during training time, we might control what feedback type we ask the human for. We might start by asking the human for demonstrations, but then move on to other types of feedback, like corrections or comparisons, to get more fine-grained information about the reward function. Since the human is only ever considering one type of feedback at a time, the conditional independence assumption makes sense.

But the assumption breaks when the human has access to multiple types of feedback at once because the types of feedback the robot can interpret influence what the human does in the first place.We note that this adaptation by the human only applies to types of behavior that the human uses to purposefully communicate with the robot, as opposed to sources of information like initial state. If the human intervenes and turns the robot off, that means one thing if this were the only feedback type available, and a whole different thing if, say, corrections were available too. In the latter case, we have more information - we know that the user chose to turn the robot off rather than provide a correction.

Thus, our key insight is that the type of feedback itself leaks information about the reward, and the RRC framework gives us a recipe for formalizing this new source: we need to uncover the set of options the human is choosing from. The human has two stages of choice: the first is the choice between feedback types, i.e corrections, language, turn-off, etc. and the second is the choice within the chosen feedback type, i.e the specific correction that the human gave. Our formalism can leverage both sources of information by defining a hierarchy of reward-rational choice.

Suppose the user has access to nn types of feedback with associated choice sets C1,…,Cn\mathcal{C}_{1},\dots,\mathcal{C}_{n}, groundings ψ1,…ψn\psi_{1},\dots\psi_{n}, and Boltzmann rationalities β1,…,βn\beta_{1},\dots,\beta_{n}. For simplicity, we assume deterministic groundings. The set of choice sets C0\mathscr{C}_{0} for the first-stage choice is {C1,…,Cn}\{\mathcal{C}_{1},\dots,\mathcal{C}_{n}\} The grounding ψ0:C→fΞ\psi_{0}:\mathcal{C}\rightarrow f_{\Xi} for the first stage choice maps a feedback type Ci\mathcal{C}_{i} to the distribution of trajectories defined by the human’s behavior and grounding in the second stage:

Finally, the probability that the human gives feedback c∗c^{*} is

The first-stage decision can be interpreted as the human metareasoning over the best type of feedback. The benefit of modeling the hierarchy is that we can cleanly separate and consider noise at both the level of metareasoning (β0\beta_{0}) and the level of execution of feedback (β1,…,βn\beta_{1},\dots,\beta_{n}). Noise at the metareasoning level models the human’s imperfection in picking the optimal type of feedback. Noise at the execution level might model the fact that the human has difficulty in physically correcting a heavy and unintuitive robot.Although we modeled rationality with respect to the reward rr that the robot should optimize, we can easily extend our formalism to capture that the person might trade-off between that and their own effort – this is especially interesting at this meta-choice level, where one type of feedback might be much more difficult and thus people might want to avoid it unless it is particularly informative.

D.2 Comparing the literal interpretation to meta-choice

We showcase the potential importance of accounting for the meta-choice in an experiment in a gridworld setting, in which an agent navigates to a goal state while avoiding lava (Figure 5, left). The reward function is a linear combination of 2 features that encode the goal and lava. The human has access to two channels of feedback: “off” and corrections. We simulate the human feedback as choosing between feedback types according to Equation 12. We manipulate three factors: 1) whether the robot is naive, i.e. only accounts for the information within the feedback type, or metareasons, i.e. accounts for the other feedback types that were available but not chosen; 2) the meta-rationality parameter β0\beta_{0} modeling human imperfection in selecting the optimal type of feedback; and 3) the location of the lava, so that the rational meta-choice changes from off to corrections. We measure regret over holdout environments.

Figure 5 (left) depicts the possible grounded trajectories for corrections and for off. For the top, off is optimal because all corrections go through lava. For the bottom, the rational meta-choice is to correct. In both cases, we find that meta-reasoning gains the learner more information, as seen in the belief (center). For the top, where the person turns it off, the robot can be more confident that lava is bad. For the bottom, the fact that the person had the off option and did not use it informs the robot about the importance of reaching the goal. This translates into lower regret (right), especially as β0\beta_{0} increases and there is more signal in the feedback type choice.

D.3 What happens when metarationality is misspecified?

In our main metareasoning experiments, we assumed that the simulated human metareasoned with β0\beta_{0} and that our algorithm somehow knew this quantity. However, in practice, we will not have access to β0\beta_{0}. This brings about an interesting question: What are the effects of inference under a misspecified β0\beta_{0}. What are the effects of overestimating or underestimating the human’s rationality?

To test this, we designed an experiment in which our simulated human provided supervision with a fixed ground truth r∗r^{*} and β0∗\beta^{*}_{0} while our algorithm performs belief updates with various β0\beta_{0} above and below β0∗\beta^{*}_{0}. The first way to measure the extent of misspecification is to measure the KL divergence between the belief induced by β0∗\beta^{*}_{0} and that induced by β0\beta_{0}.

Additionally, we wanted to measure the expected regret given a human that provides supervision with rationality β0∗\beta^{*}_{0} and the algorithm that performs belief updates with rationality β0\beta_{0}.

We plot the results in Figure 6 averaged over 50 randomly sampled reward functions and beta0∈[0.0,10.0]beta_{0}\in[0.0,10.0]. We notice that when the human does not metareason (β0∗=0.0\beta_{0}^{*}=0.0, the KL divergence in the belief distribution update is large. In comparison, with any moderate level of metareasoning β0∗=2.5,5.0,7.5\beta_{0}^{*}=2.5,5.0,7.5, the KL divergence is very low. We notice this too in the expected regret. Note that the minimum expected regret is not achieved by β0=β0∗\beta_{0}=\beta_{0}^{*}. This is because β0∗\beta_{0}^{*} is used to compute the frequency at which the human provides each type of feedback as an answer. Simply matching β0\beta_{0} with β0∗\beta_{0}^{*} doesn’t guarantee minimum expected regret (the optimal β0\beta_{0} for minimizing expected regret is a function of β0∗\beta_{0}^{*}). These experiments suggest that if we detect that the human is poor at metareasoning (low β0∗\beta_{0}^{*}), it is safer to drop the metareasoning assumption. However, if the human is displaying metareasoning, we can leverage this to improve learning.