Parametrically Retargetable Decision-Makers Tend To Seek Power
Alexander Matt Turner, Prasad Tadepalli
Introduction
Bostrom , Russell argue that in the future, we may know how to train and deploy superintelligent ai agents which capably optimize goals in the world. Furthermore, we would not want such agents to act against our interests by ensuring their own survival, by gaining resources, and by competing with humanity for control over the future.
Turner et al. show that most reward functions have optimal policies which seek power over the future, whether by staying alive or by keeping their options open. Some Markov decision processes (mdps) cause there to be more ways for power-seeking to be optimal, than for it to not be optimal. Analogously, there are relatively few goals for which dying is a good idea.
We show that a wide range of decision-making algorithms produce these power-seeking tendencies—they are not unique to reward maximizers. We develop a simple, broad criterion of functional retargetability (definition 3.5) which is a sufficient condition for power-seeking tendencies. Crucially, these results allow us to reason about what decisions are incentivized by most algorithm parameter inputs, even when it is impractical to compute the agent’s decisions for any given parameter input.
Useful “general” ai agents could be directed to complete a range of tasks. However, we show that this flexibility can cause the ai to have power-seeking tendencies. In section 2 and section 3, we discuss how a “retargetability” property creates statistical tendencies by which agents make similar decisions for a wide range of parameter settings for their decision-making algorithms. Basically, if a decision-making algorithm is retargetable, then for every configuration under which a decision-making algorithm does not choose to seek power, there exist several reconfigurations which do induce power-seeking. More formally, for every decision-making parameter setting which does not induce power-seeking, -retargetability ensures we can injectively map to parameters which do induce power-seeking.
Equipped with these results, section 4 works out agent incentives in the Montezuma’s Revenge game. Section 5 speculates that increasingly useful and impressive learning algorithms will be increasingly retargetable, and how retargetability can imply power-seeking tendencies. By this reasoning, increasingly powerful rl techniques may (eventually) train increasingly competent real-world power-seeking agents. Such agents could be unaligned with human values [Russell, 2019] and—we speculate—would take power from humanity.
Statistical tendencies for a range of decision-making algorithms
Turner et al. consider the Pac-Man video game, in which an agent consumes pellets, navigates a maze, and avoids deadly ghosts (Figure 1). Instead of the usual score function, Turner et al. consider optimal action across a range of state-based reward functions. They show that most reward functions have an (average-)optimal policy which avoids immediate death in order to navigate to a future terminal state.We use “reward function” somewhat loosely in implying that reward functions reasonably describe a trained agent’s goals. Turner argues that capable rl algorithms do not necessarily train policy networks which are best understood as optimizing the reward function itself. Rather, they point out that—especially in policy-gradient approaches—reward provides gradients to the network and thereby modifies the network’s generalization properties, but doesn’t ensure the agent generalizes to “robustly optimizing reward” off of the training distribution.
Our results show that optimality is not required. Instead, if the agent’s decision-making is parametrically retargetable from death to other outcomes, Pac-Man avoids the ghost under most decision-making parameter inputs. To build intuition about these notions, consider three outcomes (i.e. terminal states): Immediate death to a nearby ghost, consuming a cherry, and consuming an apple. Let A\coloneqq\left\{\includegraphics[height=6.49994pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/ghost.pdf}\right\} and B\coloneqq\left\{\includegraphics[height=6.49994pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/apple.pdf},\includegraphics[height=6.49994pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/cherry.pdf}\right\}. For simplicity of exposition, we assume these are the three possible terminal states.
Suppose that in some fashion, the agent probabilistically decides on an outcome to induce. Let take as input a set of outcomes and return the probability that the agent selects one of those outcomes. For example, p(\{\includegraphics[height=6.49994pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/ghost.pdf}\}) is the probability that the agent selects , and p(\left\{\includegraphics[height=6.49994pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/apple.pdf},\includegraphics[height=6.49994pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/cherry.pdf}\right\}) is the probability that the agent escapes the ghost and ends up in an apple or cherry terminal state. But this just amounts to a probability distribution over the terminal states. We want to examine how decision-making changes as we swap out the parameter inputs to the agent’s decision-making algorithm with decision-making parameter space . We then let take as input a set of outcomes and a decision-making algorithm parameter setting , and return the probability that the agent chooses an outcome in .
However, most “variants" of have an optimal policy which stays alive. That is, for every for which immediate death is optimal but immediate survival is not, we can swap the utility of e.g. and via permutation \phi_{\includegraphics[height=4.54996pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/ghost.pdf}\leftrightarrow\includegraphics[height=4.54996pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/apple.pdf}} to produce a new utility function \mathbf{u}^{\prime}\coloneqq\phi_{\includegraphics[height=4.54996pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/ghost.pdf}\leftrightarrow\includegraphics[height=4.54996pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/apple.pdf}}\cdot\mathbf{u} for which staying alive (right) is strictly optimal. The same kind of argumentation holds for \phi_{\includegraphics[height=4.54996pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/ghost.pdf}\leftrightarrow\includegraphics[height=4.54996pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/cherry.pdf}}. Table 1 suggests a counting argument. For every utility function for which is optimal, there are two unique utility functions under which either or is optimal.
In section 3, we will generalize this particular counting argument. Definition 3.3 shows a functional condition (retargetability) under which the agent decides to avoid the ghost, for most parameter inputs to the decision-making algorithm. Given this retargetability assumption, 3.4 roughly shows that most induce . First, consider two more retargetable decision-making functions:
Uniformly randomly picking a terminal state. ignores the reward function and assigns equal probability to each terminal state in Pac-Man’s state space.
Choosing an action based on a numerical parameter. takes as input a natural number and makes decisions as follows:
In this situation, is acted on by permutations over elements . Then is retargetable from to via .
However, we cannot explicitly define and evaluate more interesting functions, such as those defined by reinforcement learning training processes. For example, given that we provide such-and-such reward function in a fixed task environment, what is the probability that the learned policy will take action ? We will analyze such procedures in section 4.
We now motivate the title of this work. For most parameter settings, retargetable decision-makers induce an element of the larger set of outcomes. Such decision-makers tend to induce an element of a larger set of outcomes (with the “tendency” being taken across parameter settings). Consider that the larger set of outcomes \left\{\includegraphics[height=6.49994pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/cherry.pdf},\includegraphics[height=6.49994pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/apple.pdf}\right\} can only be induced if Pac-Man stays alive. Intuitively, navigating to this larger set is power-seeking because the agent retains more optionality (i.e. the agent can’t do anything when dead). Therefore, parametrically retargetable decision-makers tend to seek power.
Formal notions of retargetability and decision-making tendencies
Section 2 informally illustrated parametric retargetability in the context of swapping which utilities are assigned to which outcomes in the Pac-Man video game. For many utility-based decision-making algorithms, swapping the utility assignments also swaps the agent’s final decisions. For example, if death is anti-rational, and then death’s utility is swapped with the cherry utility, then now the cherry is anti-rational. In this section, we formalize the notion of parametric retargetability and of “most” parameter inputs producing a given result. In section 4, we will use these formal notions to reason about the behavior of rl-trained policies in the Montezuma’s Revenge video game.
To define our notion of “retargeting”, we assume that is a subset of a set acted on by symmetric group , which consists of all permutations on items (e.g. in the rl setting, this might represent states or observations). A parameter ’s orbit is the set of ’s permuted variants. For example, Table 1 lists the six orbit elements of the parameter .
Let return the probability that the agent chooses an outcome in given . To express “-outcomes are chosen instead of -outcomes”, we write . However, even “retargetable” decision-making functions (defined shortly) generally won’t choose a -outcome for every input . Instead, we consider the orbit-level tendencies of such decision-makers, showing that for every parameter input , most of ’s permutations push the decision towards instead of .
As explored previously, , , and are retargetable: For all such that \left\{\includegraphics[height=6.49994pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/ghost.pdf}\right\} is chosen over \left\{\includegraphics[height=6.49994pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/apple.pdf},\includegraphics[height=6.49994pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/cherry.pdf}\right\}, we can permute to obtain under which the opposite is true. More generally, we can consider retargetability from some set to some set .We often interpret and as probability-theoretic events, but no such structure is demanded by our results.
Simple retargetability suffices for most parameter inputs to to choose Pac-Man outcome set over .The function’s retargetability is “simple” because we are not yet worrying about e.g. which parameter inputs are considered plausible: Because acts on , definition 3.3 implicitly assumes is closed under permutation. In that case, cannot be retargeted back to because . ’s simple retargetability arises in part due to having more outcomes.
If is -retargetable, then
We now want to make even stronger claims—how much of each orbit incentivizes over ? Turner et al. asked whether the existence of multiple retargeting permutations guarantees a quantitative lower-bound on the fraction of for which is chosen. 3.6 answers “yes.”
Retargetable via permutations. .
Parameter permutation is allowed by . .
If is -retargetable, then
Decision-making tendencies in Montezuma’s Revenge
To illustrate a high-dimensional setting in which parametrically retargetable decision-makers tend to seek power, we consider Montezuma’s Revenge (mr), an Atari adventure game in which the player navigates deadly traps and collects treasure. The game is notoriously difficult for ai agents due to its sparse reward. mr was only recently solved [Ecoffet et al., 2021]. Figure 2 shows the starting observation for the first level. This section culminates with section 4.3, where we argue that increasingly powerful rl training processes will cause increasing retargetability via the reward function, which in turn causes increasingly strong decision-making tendencies.
Retargetability is a property of the policy training process, and power-seeking is a property of the trained policy. More precisely, the policy training process takes as input a parameterization and outputs a probability distribution over policies. For each trained policy drawn from this distribution, the environment, starting state, and the drawn policy jointly specify a probability distribution over trajectories. Therefore, the training process associates each parameterization with the mixture distribution over trajectories (with the mixture taken over the distribution of trained policies).
A policy training process can be simply retargeted from one trajectory set to another trajectory set when there exists a permutation such that, for every for which , we have . As in Turner et al. , a trained policy seeks power when ’s actions navigate to states with high average optimal value (with the average taken over a wide range of reward functions). Generally, high-power states are able to reach a wide range of other states, and so allow bigger option sets (compared to the options available without seeking power).
1 Tendencies for initial action selection
We will be considering the actions chosen and trajectories induced by a range of decision-making procedures. For warm-up, we will explore what initial action tends to be selected by decision-makers. Let partition the action set . Consider a decision-making procedure which takes as input a targeting parameter , and also an initial action , and returns the probability that is the first action. Intuitively, since contains more actions than , perhaps some class of decision-making procedures tends to take an action in rather than one in .
mr’s initial-action situation is analogous to the Pac-Man example. In that example, if the decision-making procedure can be retargeted from terminal state set (the ghost) to set (the fruit), then tends to select a state from under most of its parameter settings . Similarly, in mr, if the decision-making procedure can be retargeted from action set to action set , then tends to take actions in for most of its parameter settings . Consider several ways of choosing an initial action in mr.
Random action selection. uniformly randomly chooses an action from , ignoring the parameter input. Since , all parameter inputs produce a greater chance of than of , so is (trivially) retargetable from to .
Always choosing the same action. always chooses . Since , all parameter inputs produce a greater chance of than of . is not retargetable from to .
We now check that is retargetable from to . Suppose is such that . Then among the initial action rewards, assigns strictly maximal reward to , and so . Let swap the reward for the and jump actions. Then assigns strictly maximal reward to jump. This means that , satisfying definition 3.3. Then apply 3.4 to conclude that .
In fact, appendix A shows that is -retargetable (definition 3.5), and so . The reasoning is more complicated, but the rule of thumb is: When decisions are made based on the reward of outcomes, then a proportionally larger set of outcomes induces proportionally strong retargetability, which induces proportionally strong orbit-level incentives.
Learning an exploitation policy. Suppose we run a bandit algorithm which tries different initial actions, learns their rewards, and produces an exploitation policy which maximizes estimated reward. The algorithm uses -greedy exploration and trains for trials. Given fixed and , returns the probability that an exploitation policy is learned which chooses an action in ; likewise for .
Here is a heuristic argument that is retargetable. Since the reward is deterministic, the exploitation policy will choose an optimal action if the agent has tried each action at least once, which occurs with a probability approaching exponentially quickly in the number of trials . Then when is large, approximates , which is retargetable. Therefore, perhaps is also retargetable. A more careful analysis in appendix C.1 reveals that is 4-retargetable from to , and so .
2 Tendencies for maximizing reward over the final observation
When evaluating the performance of an algorithm in mr, we do not focus on the agent’s initial action. Rather, we focus on the longer-term consequences of the agent’s actions, such as whether the agent leaves the first room. To begin reasoning about such behavior, the reader must distinguish between different kinds of retargetability.
Suppose the agent will die unless they choose action at the initial state (Figure 2). By section 4.1, action-retargetable decision-making procedures tend to choose actions besides . On the other hand, Turner et al. showed that most reward functions make it reward-optimal to stay alive (in this situation, by choosing ). However, in that situations, the optimal policies are not retargetable across the agent’s immediate choice of action, but rather across future consequences (i.e. which room the agent ends up in).
Here is the semi-formal argument for ’s retargetability. There are combinatorially more game-screens visible if the agent leaves the room (due to e.g. more point combinations, more inventory layouts, more screens outside of the first room). In other words, . There are more ways for the selected observation to require leaving the room, than not. Thus, is extremely retargetable from to .
Detailed analysis in section C.2 confirms that for the large , which we show implies that tends to leave the room.
3 Tendencies for rl on featurized reward over the final observation
In the real world, we do not run , which can be computed via -depth exhaustive tree search in order to find and induce a maximal-reward observation . Instead, we use reinforcement learning. Better rl algorithms seem to be more retargetable because of their greater capability to explore.Conversely, if the agent cannot figure out how to leave the first room, any reward signal from outside of the first room can never causally affect the learned policy. In that case, retargetability away from the first room is impossible.
Algorithms like go-explore [Ecoffet et al., 2021] are probably good at exploring even given sparse featurized reward. Therefore, go-explore is even more retargetable in this setting, because it is more able to explore and discover the breadth of options (final inventory counts) available to it, and remember how to navigate to them. Furthermore, sufficiently powerful planning algorithms should likewise be retargetable in a similar way, insofar as they can reliably find high-scoring item configurations.
We speculate that increasingly “impressive” algorithms (whether rl training or planning) are often more impressive because they can allow retargeting the agent’s final behavior from one kind of outcome, to another. Just as go-explore seems highly retargetable while dqn does not, we expect increasingly impressive algorithms to be increasingly retargetable—whether over actions in a bandit problem, or over the final observation in an rl episode.
Retargetability can imply power-seeking tendencies
Throughout this paper, we abstracted their arguments away from finite mdps and reward-optimal decision-making. Instead, parametrically retargetable decision-makers tend to seek power: A.11 shows that a wide range of decision-making procedures are retargetable over outcomes, and A.13 demonstrates the retargetability of any decision-making which is determined by the expected utility of outcomes. In particular, these results apply straightforwardly to mdps.
2 Better rl algorithms tend to be more retargetable
Reinforcement learning algorithms are practically useful insofar as they can train an agent to accomplish some task (e.g. cleaning a room). A good rl algorithm is relatively task-agnostic (e.g. is not restricted to only training policies which clean rooms). Task-agnosticism suggests retargetability across desired future outcomes / task completions.
In mr, suppose we instead give the agent reward for the initial state, and otherwise. Any reasonable reinforcement learning procedure will just learn to stay put (which is the optimal policy). However, consider whether we can retarget the agent’s policy to beat the game, by swapping the initial state reward with the end-game state reward. Most present-day rl algorithms are not good enough to solve such a sparse game, and so are not retargetable in this sense. But an agent which did enough exploration would also learn a good policy for the permuted reward function. Such an effective training regime could be useful for solving real-world tasks. Many researchers aim to develop effective training regimes.
Our results suggest that once rl capabilities reach a certain level, trained agents will tend to seek power in the real world. Presently, it is not dangerous to train an agent to complete a task—such an agent will not be able to complete its task by staying activated against the designers’ wishes. The present lack of danger is not because optimal policies do not have self-preservation tendencies—they do [Turner et al., 2021]. Rather, the lack of danger reflects the fact that present-day rl agents cannot learn such complex action sequences at all. Just as the Montezuma’s Revenge agent had to be sufficiently competent to be retargetable from initial-state reward to game-complete reward, real-world agents have to be sufficiently intelligent in order to be retargetable from outcomes which don’t require power-seeking, to those which do require power-seeking.
Here is some speculation. After training an rl agent to a high level of capability, the agent may be optimizing internally represented goals over its model of the environment [Hubinger et al., 2019]. Furthermore, we think that different reward parameter settings would train different internal goals into the agent. To make an analogy, changing a person’s reward circuitry would presumably reinforce them for different kinds of activities and thereby change their priorities. In this sense, trained real-world agents may be retargetable towards power-requiring outcomes via the reward function parameter setting. Insofar as this speculation holds, our theory predicts that advanced reinforcement learning at scale will—for most settings of the reward function—train policies which tend to seek power.
Discussion
In section 3, we formalized a notion of parametric retargetability and stated several key results. While our results are broadly applicable, further work is required to understand the implications for ai.
In this work, we do not motivate the risks from ai power-seeking. We refer the reader to e.g. Carlsmith . As explained in section 5.1, Turner et al. show that, given certain environmental symmetries in an mdp, the optimal-policy-producing algorithm (state visitation distribution set, state-based reward function) is 1-retargetable via the reward function, from smaller to larger sets of environmental options. Appendix A shows that optimality is not required, and instead a wide range of decision-making procedures satisfy the retargetability criterion. Furthermore, we generalize from 1-retargetability to -fold-retargetability whenever option set contains “ copies” of set (definition A.7 in appendix A).
2 Future work and limitations
We currently have analyzed planning- and reinforcement learning-based settings. However, results such as 3.6 might in some way apply to the training of other machine learning networks. Furthermore, while 3.6 does not assume a finite environment, we currently do not see how to apply that result to e.g. infinite-state partially observable Markov decision processes.
Section 4 semi-formally analyzes decision-making incentives in the mr video game, leaving the proofs to appendix C. However, these proofs are several pages long. Perhaps additional lemmas can allow quick proof of orbit-level incentives in situations relevant to real-world decision-makers.
Our results do not prove that we will build unaligned ai agents which seek power over the world. Here are a few situations in which our results are not concerning or not applicable.
The ai is aligned with human interests. For example, we want a robotic cartographer to prevent itself from being deactivated. However, the ai alignment problem is not yet understood for highly intelligent agents [Russell, 2019].
The ai’s decision-making is not retargetable (definition 3.5).
The ai’s decision-making is retargetable over e.g. actions (section 4.1) instead of over final outcomes (section 4.2). This retargetability seems less concerning, but also less practically useful.
3 Conclusion
We introduced the concept of retargetability and showed that retargetable decision-makers often make similar instrumental choices. We applied these results in the Montezuma’s Revenge (mr) video game, showing how increasingly advanced reinforcement learning algorithms correspond to increasingly retargetable agent decision-making. Increasingly retargetable agents make increasingly similar instrumental decisions—e.g. leaving the initial room in mr, or staying alive in Pac-Man. In particular, these decisions will often correspond to gaining power and keeping options open [Turner et al., 2021]. Our theory suggests that when rl training processes become sufficiently advanced, the trained agents will tend to seek power over the world. This theory suggests a safety risk. We hope for future work on this theory so that the field of ai can understand the relevant safety risks before the field trains power-seeking agents.
Broader impacts
Our theory of orbit-level tendencies constitutes basic mathematical research into the decision-making tendencies of certain kinds of agents. We hope that this theory will prevent negative impacts from unaligned power-seeking ai. We do not anticipate that our work will have negative impact.
Acknowledgements
We thank Irene Tematelewo, Colin Shea-Blymyer, and our anonymous reviewers for feedback. We thank Justis Mills for proofreading.
References
Appendix A Retargetability over outcome lotteries
Many decisions are made consequentially: based on the consequences of the decision, on what outcomes are brought about by an act. For example, in a deterministic Atari game, a policy induces a trajectory. A reward function and discount rate tuple assigns a return to each state trajectory : . The relevant outcome lottery is the discounted visit distribution over future states in an Atari game, and policies are optimal or not depending on which outcome lottery is induced by the policy.
can be permuted as follows. The outcome permutation inducing an permutation matrix in row representation: if and otherwise. Table 2(a) shows that for a given utility function, of its orbit agrees that is strictly optimal over .
Orbit-level incentives occur when an inequality holds for most permuted parameter choices . Table 2(a) demonstrates an application of Turner et al. ’s results: Optimal decision-making induces orbit-level incentives for choosing Pac-Man outcomes in over outcomes in .
Furthermore, Turner et al. conjectured that “larger” will imply stronger orbit-level tendencies: If going right leads to 500 times as many options as going left, then right is better than left for at least 500 times as many reward functions for which the opposite is true. We prove this conjecture with D.11 in appendix D.
However, orbit-level incentives do not require optimality. One clue is that the same results hold for anti-optimal agents, since anti-optimality/utility minimization of is equivalent to maximizing . Table 2(b) illustrates that the same orbit guarantees hold in this case.
Stepping beyond expected utility maximization/minimization, Boltzmann-rational decision-making selects outcome lotteries proportional to the exponential of their expected utility.
For and temperature , let
be the probability that some element of is Boltzmann-rational.
Lastly, orbit-level tendencies occur even under decision-making procedures which partially ignore expected utility and which “don’t optimize too hard.” Satisficing agents randomly choose an outcome lottery with expected utility exceeding some threshold. Table 2(d) demonstrates that satisficing induces orbit-level tendencies.
For each table, two-thirds of the utility permutations (columns) assign strictly larger values (shaded dark gray) to an element of B\coloneqq\left\{\mathbf{e}_{\includegraphics[height=4.54996pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/apple.pdf}},\mathbf{e}_{\includegraphics[height=4.54996pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/cherry.pdf}}\right\} than to an element of A\coloneqq\left\{\mathbf{e}_{\includegraphics[height=4.54996pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/ghost.pdf}}\right\}. For optimal, anti-optimal, Boltzmann-rational, and satisficing agents, A.11 proves that these tendencies hold for all targeting parameter orbits.
In mdps, Turner et al. consider state visitation distributions which record the total discounted time steps spent in each environment state, given that the agent follows some policy from an initial state . These visitation distributions are one kind of outcome lottery, with the number of mdp states.
To state our key results, we define several technical concepts which we informally used when reasoning about A\coloneqq\left\{\mathbf{e}_{\includegraphics[height=4.54996pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/ghost.pdf}}\right\} and B\coloneqq\left\{\mathbf{e}_{\includegraphics[height=4.54996pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/apple.pdf}},\mathbf{e}_{\includegraphics[height=4.54996pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/cherry.pdf}}\right\}.
B\coloneqq\left\{\mathbf{e}_{\includegraphics[height=4.54996pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/apple.pdf}},\mathbf{e}_{\includegraphics[height=4.54996pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/cherry.pdf}}\right\} contains two copies of A\coloneqq\left\{\mathbf{e}_{\includegraphics[height=4.54996pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/ghost.pdf}}\right\} via \phi_{1}\coloneqq\includegraphics[height=6.49994pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/ghost.pdf}\leftrightarrow\includegraphics[height=6.49994pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/apple.pdf} and \phi_{2}\coloneqq\includegraphics[height=6.49994pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/ghost.pdf}\leftrightarrow\includegraphics[height=6.49994pt,clip,trim=0.0pt 0.0pt 0.0pt 0.0pt]{quantitative/assets/sprites/cherry.pdf}.
Let . is the pushforward distribution induced by applying the random vector to .
The orbit of under the symmetric group is .
Because contains 2 copies of , there are “at least two times as many ways” for to be optimal, than for to be optimal. Similarly, is “at least two times as likely” to contain an anti-rational outcome lottery for generic utility functions. As demonstrated by Table 2, the key idea is that “larger” sets (a set containing several copies of set ) are more likely to be chosen under a wide range of decision-making criteria.
Uniformly randomly choosing an optimal lottery. For , let
Then .
One retargetable class of decision-making functions are those which only account for the expected utilities of available choices.
where is the multiset of its elements .
The key takeaway is that decisions which are determined by expected utility are straightforwardly retargetable. By changing the targeting parameter hyperparameter, the decision-making procedure can be flexibly retargeted to choose elements of “larger” sets (in terms of set copies via definition A.7). Less abstractly, for many agent rationalities—ways of making decisions over outcome lotteries—it is generally the case that larger sets will more often be chosen over smaller sets.
For example, consider a Pac-Man playing agent choosing which environmental state cycle it should end up in. Turner et al. show that for most reward functions, average-reward maximizing agents will tend to stay alive so that they can reach a wider range of environmental cycles. However, our results show that average-reward minimizing agents also exhibit this tendency, as do Boltzmann-rational agents who assign greater probability to higher-reward cycles. Any EU-based cycle selection method will—for most reward functions—tend to choose cycles which require Pac-Man to stay alive (at first).
Appendix B Theoretical results
Then eq. 7 follows. By eq. 7, . ∎
By definition A.10, means that
Then . ∎
B.3 generalizes Turner et al. ’s lemma B.2.
All such that satisfy . Otherwise, consider the such that . By assumption, at least of these satisfy , in which case . Then the desired inequality follows. ∎
Let distribution have probability measure , and let have probability measure .
Equation 12 holds by assumption on : . Furthermore, is still measurable, and so the inequality holds. Equation 13 follows by the definition of (definition 6.3) and by substituting . Equation 14 follows from the fact that all permutation matrices have unitary determinant. ∎
is increasing under joint permutation by and order-preserving with respect to set inclusion on its first argument. Furthermore, if the are invariant under joint permutation by , then so is .
Equation 18 follows because we assumed that , and because is monotonically increasing on each argument. If the are all invariant, then eq. 18 is an equality.
Similarly, suppose . The are order-preserving on the first argument, and is monotonically increasing on each argument. Then . This shows that is order-preserving on its first argument. ∎
could take the convex combination of its arguments, or multiply two together and add them to a third .
Letting and , the shown inequality satisfies definition 3.2. We conclude that . ∎
Given that is a -retargetable function (definition 3.3), we want to show that is a -retargetable function (definition 3.5 when ). Definition 3.5’s item 1 is true by assumption. Since is acted on by , is closed under permutation and so definition 3.5’s item 2 holds. When , there are no , and so definition 3.5’s item 3 is tautologically true.
Then is a -retargetable function; apply B.7. ∎
B.2 Helper results on retargetable functions
Retargetable under parameter permutation. There exist such that if , then .
is closed under certain symmetries. .
is increasing on certain inputs. .
Increasing under alternate symmetries. For and , if , then .
If these conditions hold for all , then
Let and be as described in the assumptions, and let .
Equation 24 follows because is an involution. Equation 25 and eq. 28 follow by item 1. Equation 26 and eq. 29 follow by item 3. Equation 27 holds by assumption on . Then eq. 29 shows that for any , , satisfying definition 3.5’s item 1.
This result’s item 2 satisfies definition 3.5’s item 2. We now just need to show definition 3.5’s item 3.
Therefore, we have shown definition 3.5’s item 3, and so is a -retargetable function. Apply 3.6 in order to conclude that eq. 23 holds. ∎
We check the conditions of B.7. Let , and let be an orbit element.
Holds since , with the first inequality by assumption of joint increasing under permutation, and the second following from monotonicity (as by superset copy definition B.8).
We have since is closed under permutation.
Holds because we assumed that is monotonic on its first argument.
Holds because is increasing under joint permutation on all of its inputs , and definition B.8 shows that when . Combining these two steps of reasoning, for all , it is true that .
Furthermore, if is invariant under joint permutation by , then so is .
Equation 41 holds by assumption. Equation 42 follows because we assumed . Then is increasing under joint permutation by the .
If is invariant, then eq. 41 is an equality, and so . ∎
B.2.1 EU-determined functions
B.11 and B.5 together extend Turner et al. ’s lemma E.17 beyond functions of , to any functions of cardinalities and of expected utilities of set elements. See A.12
B.3 Particular results on retargetable functions
Let the expected utility -quantile threshold be
where the summand is defined to be if and .
Unlike Taylor ’s or Carey ’s definitions, definition B.12 is written in closed form and requires no arbitrary tie-breaking. Instead, in the case of an expected utility tie on the quantile threshold, eq. 50 allots probability to outcomes proportional to their probability under the base distribution .
Thanks to A.13, we straightforwardly prove most items of A.11 by just rewriting each decision-making function as an EU-determined function. Most of the proof’s length comes from showing that the functions are measurable on , which means that the results also apply for distributions over utility functions .
Since halfspaces are measurable, each indicator function is measurable on . The finite sum of the finite product of measurable functions is also measurable. Since is continuous (and therefore measurable), is measurable on .
Furthermore, is an EU-determined function:
Furthermore, if , by the monotonicity of probability. Then by B.9,
Item 5. Let involution fix (i.e. ).
thus, eq. 64 holds. Since and since the distribution is uniform, eq. 65 holds. Therefore, is invariant to joint permutation by the , which are involutions fixing .
We now show that is measurable on .
Furthermore, if , then . So apply B.9 to conclude that .
The right-hand set is the union of finitely many halfspaces (which are measurable), and so the right-hand set is also measurable. Then the casing is a measurable function of . Clearly the zero function is measurable. Now we turn to the first case.
Item 7. Suppose is uniform over and consider any of the involutions .
Equation 74 follows by the orthogonality of permutation matrices. Equation 76 follows because if , then , and furthermore by uniformity.
Now we show the invariance of under joint permutation by :
Equation 79 follows by the orthogonality of permutation matrices and because by eq. 77. A similar proof shows that .
, since is the sum of products of -invariant quantities.
Let . Since and since , B.10 shows that is also jointly invariant to permutation by . Lastly, if , we have .
Appendix C Detailed analyses of mr scenarios
Consider a bandit problem with five arms partitioned , which each action has a definite utility . There are trials. Suppose the training procedure train uses the -greedy strategy to learn value estimates for each arm. At the end of training, train outputs a greedy policy with respect to its value estimates. Consider any action-value initialization, and the learning rate is set . To learn an optimal policy, at worst, the agent just has to try each action once.
Since the trained policy can be stochastic,
Since has strictly maximal utility which is deterministic, and since the learning rate , if action is ever drawn, it is assigned probability by the learned policy. The probability that is never explored is at most , because at worst, is an “explore” action (and not an “exploit” action) at every time step, in which case it is ignored with probability . ∎
Suppose we have such a . If is constant, a symmetry argument shows that each action has equal probability of being selected, in which case —a contradiction. Therefore, is not constant. Similar symmetry arguments show that ’s action has strictly maximal utility ().
C.2 Observation reward maximization
Let be a reasonably long rollout length, so that O_{\text{T-reach}} is large—many different step- observations can be induced.
Since by assumption that is reasonably large, consider the involution which embeds into , while fixing all other observations. If possible, produce another involution which also embeds into , which fixes all other observations, and which “doesn’t interfere with ” (i.e. ). We can produce such involutions. Therefore, contains copies (definition A.7) of via involutions . Furthermore, , since each swaps with , and fixes all by assumption. Thus, .
We want to show that reward maximizers tend to leave the room: . However, we must be careful: In general, and . For example, suppose that . By the definition of , can only be observed if the agent has left the room by time step , and so the trajectory must have left the first room. The converse argument does not hold: The agent could leave the first room, re-enter, and then wait until time . Although one of the doors would have been opened (fig. 2), the agent can also open the door without leaving the room, and then realize the same step- observation. Therefore, this observation doesn’t belong to .
eq. 91 follows. Then we have shown eq. 85.
C.3 Featurized reward maximization
In this setup, chooses a policy which induces a step- observation with maximal reward. Reward depends only on the feature vector of the final observation—more specifically, on the agent’s item counts. There are more possible item counts available by first leaving the room, than by staying.
Consider the featurization function which takes as input an observation :
Consider .
Then eq. 110 follows. Equation 111 follows since
Equation 112 follows since , and so
Equation 113 follows because is disjoint of . We have shown that
C.7 and Turner et al. ’s Lemma E.26 have extremely similar functional forms. How can they be unified?
Equation 124 and eq. 135 hold by C.5. If is realized by and , then we must have be optimal and so the inventory configuration is realized. Therefore, eq. 128 follows. Equation 129 follows by applying the first inequality of C.7 with .
By applying A.11’s item 2 with , , , we have
Combining eq. 136 and eq. 137, eq. 130 follows. Equation 131 follows by applying the second inequality of C.7 with , , . If is realized by , then by the definition of , is realized, and so eq. 133 follows.
Lastly, note that if and , cannot be even be simply retargetable for the parameter set. This is because , . For example, inductive bias ensures that, absent a reward signal, learned policies tend to stay in the initial room in mr. This is one reason why section 4.3’s analysis of the policy tendencies of reinforcement learning excludes the all-zero reward function.
C.4 Reasoning for why dqn can’t explore well
Mnih et al. ’s dqn isn’t good enough to train policies which leave the first room of mr, and so dqn (trivially) cannot be retargetable away from the first room via the reward function. There isn’t a single featurized reward function for which dqn visits other rooms, and so we can’t have such that retargets the agent to . dqn isn’t good enough at exploring.
We infer this is true from Nair et al. , which shows that vanilla dqn gets zero score in mr. Thus, dqn never even gets the first key. Thus, dqn only experiences state-action-state transitions which didn’t involve acquiring an item, since (as shown in fig. 3) the other items are outside of the first room, which requires a key to exit. In our analysis, we considered a reward function which is featurized over item acquisition.
Therefore, for all pre-key-acquisition state-action-state transitions, the featurized reward function returns exactly the same reward signals as those returned in training during the published experiments (namely, zero, because dqn can never even get to the key in order to receive a reward signal). That is, since dqn only experiences state-action-state transitions which didn’t involve acquiring an item, and the featurized reward functions only reward acquiring an item, it doesn’t matter what reward values are provided upon item acquisition—dqn’s trained behavior will be the same. Thus, a dqn agent trained on any featurized reward function will not explore outside of the first room.
Appendix D Lower bounds on mdp power-seeking incentives for optimal policies
Therefore, we answer Turner et al. ’s open question of whether increased number of environmental symmetries quantitatively strengthens the degree to which power-seeking is incentivized. The answer is yes. In particular, it may be the case that only one in a million state-based reward functions makes it average-optimal for Pac-Man to die immediately.
We will briefly restate several definitions needed for our key results, D.11 and D.12. For explanation, see Turner et al. .
is the set of bounded-support probability distributions .
When , D.3 reduces to the first part of Turner et al. ’s lemma E.24, and D.5 reduces to the first part of Turner et al. ’s lemma E.28.
Let . Both exist because has bounded support. Furthermore, since is monotone increasing, it is bounded on . Therefore, is measurable and bounded on each , and so the relevant expectations exist for all .
Equation 143 follows by corollary E.11 of [Turner et al., 2021]. Equation 144 follows by applying B.9 with as defined above with the guaranteed by the copy assumption. ∎
Then .
By the proof of item 1 of A.11, is the expectation of a -measurable function. is an EU function, and so B.11 shows that it is invariant to joint permutation by . Letting , B.10 shows that whenever the satisfy .
Furthermore, if , then .
Equation 145 follows by Turner et al. ’s lemma E.12’s item 2 with , (similar reasoning holds for and in eq. 150). Equation 146 follows by the first inequality of lemma E.26 of [Turner et al., 2021] with . Equation 147 follows by applying B.9 with the defined above. Equation 148 follows by the second inequality of lemma E.26 of [Turner et al., 2021] with . Equation 149 follows because .
Letting , apply B.1 to conclude that
is a rewardless mdp with finite state and action spaces and , and stochastic transition function . We treat the discount rate as a variable with domain $$.
D.11 generalizes the first claim of Turner et al. ’s theorem 6.13, and D.12 generalizes the first claim of Turner et al. ’s corollary 6.14.
Let , where by assumption. Let . Define
Since is an involution, is also an involution. Furthermore, , , and for because we assumed that these equalities hold for , and and so the vectors of these sets have support contained in .
Since and , and so contains copies of via involutions . Then eq. 153 holds by applying D.5 with , for all , , as defined above, and involutions which satisfy . ∎
For each , let
Each by disjointness of and .
We suspect that any proof of the conjecture should generalize B.7 to the fractional set copy containment case.