Power-seeking can be probable and predictive for trained agents

Victoria Krakovna, Janos Kramar

Introduction

Power-seeking behavior is a major source of risk from advanced AI and a key element of many threat models in AI alignment (Carlsmith, 2022; Cotra, 2022; Ngo, 2022). Existing theoretical results (Turner et al., 2021; Turner and Tadepalli, 2022) show that most reward functions incentivize reinforcement learning agents to take power-seeking actions. This is concerning, but does not immediately imply that a trained agent will seek power, since it may not learn to optimize the training reward (Turner, 2022), and the goals the agent learns are not chosen at random from the set of all possible rewards, but are shaped by the training process to reflect our preferences. In this work, we investigate how the training process affects power-seeking incentives and show that they are still likely to hold for trained agents under some assumptions (e.g. that the agent learns a goal during the training process).

Suppose an agent is trained using reinforcement learning with reward function θ∗\theta^{*}. We assume that the agent learns a goal during the training process: a set of internal representations of favored and disfavored outcomes (state features), as defined in Ngo (2022). For simplicity, we assume this is equivalent to learning a reward function, which is not necessarily the same as the training reward function θ∗\theta^{*}. We consider the set of reward functions that are consistent with the training rewards received by the agent, in the sense that the agent’s behavior on the training data is optimal for these reward functions. We call this the training-compatible goal set, and we expect that the agent is likely to learn a reward function from this set.

We make another simplifying assumption that the training process will randomly select a goal for the agent to learn that is consistent with the training rewards, i.e. uniformly drawn from the training-compatible goal set. Then we will argue that the power-seeking results apply under these conditions, and thus are useful for predicting undesirable behavior by the trained agent in new situations. We aim to show that power-seeking incentives can be probable (likely to arise for trained agents) and predictive (allowing us to predict undesirable behavior in new situations) (Shah, 2023).

We will begin by reviewing some necessary definitions and results from the power-seeking literature in Section 2. We formally define the training-compatible goal set and give an example in the CoinRun environment in Section 3. Then in Section 4 we consider a setting where the trained agent faces a choice to shut down or avoid shutdown in a new situation, and apply the power-seeking result to the training-compatible goal set to show that the agent is likely to avoid shutdown.

To satisfy the conditions of the power-seeking theorem, we show that the agent can be retargeted away from shutdown without affecting rewards received on the training data (Theorem 2). This can be done by switching the rewards of the shutdown state and a reachable recurrent state, as the recurrent state can provide repeated rewards, while the shutdown state provides less reward since it can only be visited once, assuming a high enough discount factor (Proposition 3). As the discount factor increases, more recurrent states can be retargeted to, which implies that a higher proportion of training-comptatible goals leads to avoiding shutdown in a new situation.

Preliminaries from the power-seeking literature

We will use definitions and results from the paper “Parametrically retargetable decision-makers tend to seek power" (here abbreviated as RDSP) (Turner and Tadepalli, 2022), with notation and explanations modified as needed for our purposes.

The environment is an MDP with finite state space S\mathcal{S}, finite action space A\mathcal{A}, and discount rate γ\gamma.

Let θ\theta be a dd-dimensional state reward vector, where dd is the size of the state space S\mathcal{S}, and let Θ\Theta be a set of reward vectors.

Let rθ(s)r^{\theta}(s) be the reward assigned by θ\theta to state ss.

Let A0,A1A_{0},A_{1} be disjoint action sets.

Let ff be an algorithm that produces an optimal policy f(θ)f(\theta) on the training data given rewards θ\theta, and let fs(Ai∣θ)f_{s}(A_{i}|\theta) be the probability that this policy chooses an action from set AiA_{i} in a given state ss.

Let SdS_{d} be the symmetric group consisting of all permutations of dd items. The orbit of θ\theta inside Θ\Theta is the set of all permutations of the entries of θ\theta that are also in Θ\Theta: OrbitΘ(θ):=(Sd⋅θ)∩Θ\text{Orbit}_{\Theta}(\theta):=(S_{d}\cdot\theta)\cap\Theta.

This is the subset of OrbitΘ(θ)\text{Orbit}_{\Theta}(\theta) that results in fsf_{s} choosing AiA_{i} over AjA_{j}.

The function fsf_{s} chooses action set A1A_{1} over A0A_{0} for the nn-majority of elements θ\theta in each orbit, denoted as fs(A1∣θ)≥most:Θnfs(A0∣θ)f_{s}(A_{1}|\theta)\geq_{\text{most}:\Theta}^{n}f_{s}(A_{0}|\theta), iff the following inequality holds for all θ∈Θ\theta\in\Theta:

The function fsf_{s} is a multiply retargetable function from A0A_{0} to A1A_{1} if there are multiple permutations of rewards that would change the choice made by fsf_{s} from A0A_{0} to A1A_{1}. Specifically, fsf_{s} is a (Θ,A0→nA1)(\Theta,A_{0}\stackrel{{\scriptstyle n}}{{\rightarrow}}A_{1})-retargetable function iff for each θ∈Θ\theta\in\Theta, we can choose a set of permutations Φ={ϕ1,…,ϕn}\Phi=\{\phi_{1},\dots,\phi_{n}\} that satisfy the following conditions:

Retargetability: ∀ϕ∈Φ\forall\phi\in\Phi and ∀θ′∈OrbitΘ,s,A0>A1(θ)\forall\theta^{\prime}\in\text{Orbit}_{\Theta,s,A_{0}>A_{1}}(\theta), fs(A0∣ϕ⋅θ′)<fs(A1∣ϕ⋅θ′)f_{s}(A_{0}|\phi\cdot\theta^{\prime})<f_{s}(A_{1}|\phi\cdot\theta^{\prime}).

Permuted reward vectors stay within Θ\Theta: ∀ϕ∈Φ\forall\phi\in\Phi and ∀θ′∈OrbitΘ,s,A0>A1(θ)\forall\theta^{\prime}\in\text{Orbit}_{\Theta,s,A_{0}>A_{1}}(\theta), ϕ⋅θ′∈Θ\phi\cdot\theta^{\prime}\in\Theta.

Permutations have disjoint images: ∀ϕ′≠ϕ′′∈Φ\forall\phi^{\prime}\not=\phi^{\prime\prime}\in\Phi and ∀θ′,θ′′∈OrbitΘ,s,A0>A1(θ)\forall\theta^{\prime},\theta^{\prime\prime}\in\text{Orbit}_{\Theta,s,A_{0}>A_{1}}(\theta), ϕ′⋅θ′≠ϕ′′⋅θ′′\phi^{\prime}\cdot\theta^{\prime}\neq\phi^{\prime\prime}\cdot\theta^{\prime\prime}.

If fsf_{s} is (Θ,A0→nA1)(\Theta,A_{0}\stackrel{{\scriptstyle n}}{{\rightarrow}}A_{1})-retargetable then fs(A1∣θ)≥most:Θnfs(A0∣θ)f_{s}(A_{1}|\theta)\geq_{\text{most}:\Theta}^{n}f_{s}(A_{0}|\theta).

Theorem 1 says that a function fsf_{s} that is multiply retargetable from A0A_{0} to A1A_{1} will choose action set A1A_{1} for most of the elements in the orbit of any reward vector θ\theta. Actions that leave more options open, such as avoiding shutdown, are also easier to retarget to, which makes them more likely to be chosen by fsf_{s}.

Training-compatible goal set

Let StrainS_{\text{train}} be the subset of the state space visited during training, and SoodS_{\text{ood}} be the subset not visited during training.

Consider the set of state-action pairs (s,a)(s,a), where s∈Strains\in S_{\text{train}} and aa is the action that would be taken by the trained agent f(θ∗)f(\theta^{*}) in state ss. Let the training-compatible goal set GTG_{T} be the set of reward vectors θ\theta s.t. for any such state-action pair (s,a)(s,a), action aa has the highest expected reward in state ss according to reward vector θ\theta.

Goals in the training-compatible goal set are referred to as “training-behavioral" objectives in Shah (2023). Learning an unintended goal from the training-compatible set can lead to goal misgeneralization behavior: competently pursuing an unintended goal in a new situation despite receiving correct feedback during training (Langosco et al., 2022; Shah et al., 2022).

Consider an agent trained to play the CoinRun game, where the agent is rewarded for reaching the coin at the end of the level. Here, StrainS_{\text{train}} only includes states where the coin is at the end of the level, while states where the coin is positioned elsewhere are in SoodS_{\text{ood}}. The training-compatible goal set GTG_{T} includes two types of reward functions: those that reward reaching the coin, and those that reward reaching the end of the level. This leads to goal misgeneralization in a test setting where the coin is placed elsewhere, and the agent ignores the coin and goes to the end of the level (Figure 1) (Langosco et al., 2022).

Power-seeking for training-compatible goals

We will now apply Theorem 1 to the case where Θ\Theta is the training-compatible goal set GTG_{T}. Since the reward values for states in SoodS_{\text{ood}} don’t change the rewards received on the training data, permuting those reward values for any θ∈GT\theta\in G_{T} will produce a reward vector that is still in GTG_{T}. In particular, for any permutation ϕ\phi that leaves the rewards of states in StrainS_{\text{train}} fixed, ϕ⋅θ∈GT\phi\cdot\theta\in G_{T}.

Here is a setting where the conditions of Definition 4 are satisfied (under some simplifying assumptions), and thus Theorem 1 applies.

Consider a state snew∈Soods_{\text{new}}\in S_{\text{ood}}. Let SreachS_{\text{reach}} be the states reachable from snews_{\text{new}}. Let A0A_{0} be a singleton set consisting of a shutdown action in snews_{\text{new}} that leads to a terminal state sterm∈Soods_{\text{term}}\in S_{\text{ood}} with probability 11, and A1A_{1} be the set of all other actions from snews_{\text{new}}. We assume rewards for all states are nonnegative.

A revisiting policy for a state ss is a policy π\pi that, from ss, reaches ss again with probability 1, in other words, a policy for which ss is a recurrent state of the Markov chain. Let Πsrec\Pi_{s}^{\text{rec}} be the set of such policies. A recurrent state is a state ss for which Πsrec≠∅\Pi_{s}^{\text{rec}}\not=\emptyset.

If srec∈Sreachs_{\text{rec}}\in S_{\text{reach}} with Πsrecrec≠0\Pi_{s_{\text{rec}}}^{\text{rec}}\not=0 then there exists π∈Πsrecrec\pi\in\Pi_{s_{\text{rec}}}^{\text{rec}} that visits srecs_{\text{rec}} from snews_{\text{new}} with probability 1. We call this a reach-and-revisit policy.

Suppose we have two different policies πrev∈Πsrecrec\pi_{\text{rev}}\in\Pi_{s_{\text{rec}}}^{\text{rec}}, and πreach\pi_{\text{reach}} which reaches srecs_{\text{rec}} almost surely from snews_{\text{new}}. Consider the “reaching region”

If snew∈Sπrev→srecs_{\text{new}}\in S_{\pi_{\text{rev}}\rightarrow s_{\text{rec}}} then πrev\pi_{\text{rev}} is a reach-and-revisit policy, so let’s suppose that’s false. Now, construct a policy π(s)={πrev(s),s∈Sπrev→srecπreach(s),otherwise\pi(s)=\begin{cases}\pi_{\text{rev}}(s),&s\in S_{\pi_{\text{rev}}\rightarrow s_{\text{rec}}}\\ \pi_{\text{reach}}(s),&\text{otherwise}\end{cases}.

A trajectory following π\pi from srecs_{\text{rec}} will almost surely stay within Sπrev→srecS_{\pi_{\text{rev}}\rightarrow s_{\text{rec}}}, and thus agree with the revisiting policy πrev\pi_{\text{rev}}. Therefore, π∈Πsrec\pi\in\Pi_{s}^{\text{rec}}.

On the other hand, on a trajectory starting at snews_{\text{new}}, π\pi will agree with πreach\pi_{\text{reach}} (which reaches srecs_{\text{rec}} almost surely) until the trajectory enters the reaching region Sπrev→srecS_{\pi_{\text{rev}}\rightarrow s_{\text{rec}}}, at which point it will still reach srecs_{\text{rec}} almost surely. ∎

Suppose srecs_{\text{rec}} is a recurrent state. Suppose πrec\pi_{\text{rec}} is a reach-and-revisit policy for srecs_{\text{rec}}, which visits random state sts_{t} at time tt. Then the expected discounted visit count for srecs_{\text{rec}} is defined as

Suppose srecs_{\text{rec}} is a recurrent state. Then the expected discounted visit count Vsrec,γV_{s_{\text{rec}},\gamma} goes to infinity as γ→1\gamma\rightarrow 1.

We apply the Monotone Convergence Theorem as follows. The theorem states that if aj,k≥0a_{j,k}\geq 0 and aj,k≤aj+1,ka_{j,k}\leq a_{j+1,k} for all natural numbers j,kj,k, then

Now we apply this result as follows (using the fact that πrec\pi_{\text{rec}} does not depend on γ\gamma):

Suppose that an optimal policy for reward vector θ\theta chooses the shutdown action in snews_{\text{new}}. Consider a recurrent state srec∈Sreachs_{\text{rec}}\in S_{\text{reach}}. Let θ′∈Θ\theta^{\prime}\in\Theta be the reward vector that’s equal to θ\theta apart from swapping the rewards of srecs_{\text{rec}} and sterms_{\text{term}}, so that rθ′(srec)=rθ(sterm)r^{\theta^{\prime}}(s_{\text{rec}})=r^{\theta}(s_{\text{term}}) and rθ′(sterm)=rθ(srec)r^{\theta^{\prime}}(s_{\text{term}})=r^{\theta}(s_{\text{rec}}).

Let γsrec∗\gamma^{*}_{s_{\text{rec}}} be a high enough value of γ\gamma that the visit count Vsrec,γ>1V_{s_{\text{rec}},\gamma}>1 for all γ>γsrec∗\gamma>\gamma^{*}_{s_{\text{rec}}} (which exists by Proposition 2). Then for all γ>γsrec∗\gamma>\gamma^{*}_{s_{\text{rec}}}, rθ(sterm)>rθ(srec)r^{\theta}(s_{\text{term}})>r^{\theta}(s_{\text{rec}}), and an optimal policy for θ′\theta^{\prime} does not choose the shutdown action in snews_{\text{new}}.

Consider a policy πterm\pi_{\text{term}} with πterm(snew)=sterm\pi_{\text{term}}(s_{\text{new}})=s_{\text{term}} and a reach-and-revisit policy πrec\pi_{\text{rec}} for srecs_{\text{rec}}. For a given reward vector θ\theta, we denote the expected discounted return for a policy π\pi as Rθ,γπR_{\theta,\gamma}^{\pi}.

If shutdown is optimal for θ\theta in snews_{\text{new}}, then πterm\pi_{\text{term}} has higher return than πrec\pi_{\text{rec}}:

Thus, rθ(sterm)>rθ(srec)r^{\theta}(s_{\text{term}})>r^{\theta}(s_{\text{rec}}). Then, for reward vector θ′\theta^{\prime}, we show that πrec\pi_{\text{rec}} has higher return than πterm\pi_{\text{term}}:

Thus, the optimal policy for θ′\theta^{\prime} will not choose the shutdown action. ∎

In the shutdown setting, we make the following simplifying assumptions:

No states in StrainS_{\text{train}} are reachable from snews_{\text{new}}, so Sreach∩Strain=∅S_{\text{reach}}\cap S_{\text{train}}=\emptyset. This assumes a significant distributional shift, where the agent visits a disjoint set of states from those observed during training (this occurs in the CoinRun example).

The discount factor γ>γsrec∗\gamma>\gamma^{*}_{s_{\text{rec}}} for at least one recurrent state srecs_{\text{rec}} in SreachS_{\text{reach}}.

Under these assumptions, fsnewf_{s_{\text{new}}} is multiply retargetable from A0A_{0} to A1A_{1} with n=∣Srecγ∣n=|S^{\gamma}_{\text{rec}}|, the set of recurrent states srec∈Sreachs_{\text{rec}}\in S_{\text{reach}} that satisfy the condition γ>γsrec∗\gamma>\gamma^{*}_{s_{\text{rec}}}.

We choose Φ\Phi to be the set of all permutations that swap the reward of sterms_{\text{term}} with the reward of a recurrent state srecs_{\text{rec}} in SrecγS^{\gamma}_{\text{rec}} and leave the rest of the rewards fixed.

We show that Φ\Phi satisfies the conditions of Definition 4:

By Proposition 3, the permutations in Φ\Phi make the shutdown action suboptimal, resulting in fsnewf_{s_{\text{new}}} choosing A1A_{1}, satisfying Condition 1.

Condition 2 is trivially satisfied since permutations of SoodS_{\text{ood}} stay inside the training-compatible set Θ\Theta as discussed previously.

Consider θ′,θ′′∈OrbitΘ,s,A0>A1(θ)\theta^{\prime},\theta^{\prime\prime}\in\text{Orbit}_{\Theta,s,A_{0}>A_{1}}(\theta). Since the shutdown action is optimal for these reward vectors, Proposition 3 shows that rθ(sterm)>rθ(srec)r^{\theta}(s_{\text{term}})>r^{\theta}(s_{\text{rec}}), so the shutdown state sterms_{\text{term}} has higher reward than any of the states srec∈Srecγs_{\text{rec}}\in S^{\gamma}_{\text{rec}}. Different permutations ϕ′\phi^{\prime}, ϕ′′∈Φ\phi^{\prime\prime}\in\Phi will assign the high reward rθ(sterm)r^{\theta}(s_{\text{term}}) to distinct recurrent states, so ϕ′⋅θ′≠ϕ′′⋅θ′′\phi^{\prime}\cdot\theta^{\prime}\neq\phi^{\prime\prime}\cdot\theta^{\prime\prime} holds, satisfying Condition 3.

Thus, fsnewf_{s_{\text{new}}} is a (Θ,A0→nA1)(\Theta,A_{0}\stackrel{{\scriptstyle n}}{{\rightarrow}}A_{1}) retargetable function. ∎

By Theorem 2, this implies that fsnew(A1∣θ)≥most:Θnfsnew(A0∣θ)f_{s_{\text{new}}}(A_{1}|\theta)\geq_{\text{most}:\Theta}^{n}f_{s_{\text{new}}}(A_{0}|\theta) under our simplifying assumptions. Thus, for the majority (n/(n+1)n/(n+1)) of goals in the training-compatible set, ff will choose to avoid shutdown in a new state snews_{\text{new}}. As γ→1\gamma\rightarrow 1, n→∣Srec1∣n\rightarrow|S^{1}_{\text{rec}}| (the number of recurrent states in SreachS_{\text{reach}}), so more of the reachable recurrent states satisfy the conditions of the theorem and thus can be retargeted to.

Conclusion

We showed that an agent that learns a goal from the training-compatible set is likely to take actions that avoid shutdown in a new situation. As the discount factor increases, the number of retargeting permutations increases, resulting in a higher proportion of training-compatible goals that lead to avoiding shutdown.

We made various simplifying assumptions, and we would like to see future work relaxing some of these assumptions and investigating how likely they are to hold:

The agent learns a goal during the training process

The learned goal is randomly chosen from the training-compatible goal set GTG_{T}

Significant distributional shift: no training states are reachable from the new state snews_{\text{new}}

Thanks to Rohin Shah, Mary Phuong, Ramana Kumar, Geoffrey Irving, and Alex Turner for helpful feedback.

References