Power-seeking can be probable and predictive for trained agents

Link post

Power-seeking is a major source of risk from advanced AI and a key element of most threat models in alignment. Some theoretical results show that most reward functions incentivize reinforcement learning agents to take power-seeking actions. This is concerning, but does not immediately imply that the agents we train will seek power, since the goals they learn are not chosen at random from the set of all possible rewards, but are shaped by the training process to reflect our preferences. In this work, we investigate how the training process affects power-seeking incentives and show that they are still likely to hold for trained agents under some assumptions (e.g. that the agent learns a goal during the training process).

Suppose an agent is trained using reinforcement learning with reward function $θ^{*}$ . We assume that the agent learns a goal during the training process: some form of implicit internal representation of desired state features or concepts. For simplicity, we assume this is equivalent to learning a reward function, which is not necessarily the same as the training reward function $θ^{*}$ . We consider the set of reward functions that are consistent with the training rewards received by the agent, in the sense that agent’s behavior on the training data is optimal for these reward functions. We call this the training-compatible goal set, and we expect that the agent is most likely to learn a reward function from this set.

We make another simplifying assumption that the training process will randomly select a goal for the agent to learn that is consistent with the training rewards, i.e. uniformly drawn from the training-compatible goal set. Then we will argue that the power-seeking results apply under these conditions, and thus are useful for predicting undesirable behavior by the trained agent in new situations. We aim to show that power-seeking incentives are probable and predictive: likely to arise for trained agents and useful for predicting undesirable behavior in new situations.

We will begin by reviewing some necessary definitions and results from the power-seeking literature. We formally define the training-compatible goal set (Definition 7) and give an example in the CoinRun environment. Then we consider a setting where the trained agent faces a choice to shut down or avoid shutdown in a new situation, and apply the power-seeking result to the training-compatible goal set to show that the agent is likely to avoid shutdown.

To satisfy the conditions of the power-seeking theorem (Theorem 1), we show that the agent can be retargeted away from shutdown without affecting rewards received on the training data (Theorem 2). This can be done by switching the rewards of the shutdown state and a reachable recurrent state, as the recurrent state can provide repeated rewards, while the shutdown state provides less reward since it can only be visited once, assuming a high enough discount factor (Proposition 3). As the discount factor increases, more recurrent states can be retargeted to, which implies that a higher proportion of training-comptatible goals leads to avoiding shutdown in a new situation.

Preliminaries from the power-seeking literature

We will rely on the following definitions and results from the paper Parametrically retargetable decision-makers tend to seek power (here abbreviated as RDSP), with notation and explanations modified as needed for our purposes.

Notation and assumptions

The environment is an MDP with finite state space $S$ , finite action space $A$ , and discount rate $γ$ .
Let $θ$ be a $d$ -dimensional state reward vector, where $d$ is the size of the state space $S$ and let $Θ$ be a set of reward vectors.
Let $r^{θ} (s)$ be the reward assigned by $θ$ to state $s$ .
Let $A_{0}, A_{1}$ be disjoint action sets.
Let $f$ be an algorithm that produces an optimal policy $f (θ)$ on the training data given rewards $θ$ , and let $f_{s} (A_{i} | θ)$ be the probability that this policy chooses an action from set $A_{i}$ in a given state $s$ .

Definition 1: Orbit of a reward vector (Def 3.1 in RDSP)

Let $S_{d}$ be the symmetric group consisting of all permutations of d items.

The orbit of $θ$ inside $Θ$ is the set of all permutations of the entries of $θ$ that are also in $Θ$ : ${Orbit}_{Θ} (θ) := (S_{d} \cdot θ) \cap Θ$ .

Definition 2: Orbit subset where an action set is preferred (from Def 3.5 in RDSP)

Let ${Orbit}_{Θ, s, A_{i} > A_{j}} (θ) := {θ^{'} \in {Orbit}_{Θ} (θ) | f_{s} (A_{i} | θ^{'}) > f_{s} (A_{j} | θ^{'})}$ . This is the subset of ${Orbit}_{Θ} (θ)$ that results in $f_{s}$ choosing $A_{i}$ over $A_{j}$ .

Definition 3: Preference for an action set $A_{1}$ (Def 3.2 in RDSP)

The function $f_{s}$ chooses action set $A_{1}$ over $A_{0}$ for the $n$ -majority of elements $θ$ in each orbit, denoted as $f_{s} (A_{1} | θ) \geq_{most : Θ}^{n} f_{s} (A_{0} | θ)$ , iff the following inequality holds for all $θ \in Θ$ : $∣ ∣ {Orbit}_{Θ, s, A_{1} > A_{0}} (θ) ∣ ∣ \geq n ∣ ∣ {Orbit}_{Θ, s, A_{0} > A_{1}} (θ) ∣ ∣$ .

Definition 4: Multiply retargetable function from $A_{0}$ to $A_{1}$ (Def 3.5 in RDSP)

The function $f_{s}$ is a multiply retargetable function from $A_{0}$ to $A_{1}$ if there are multiple permutations of rewards that would change the choice made by $f_{s}$ from $A_{0}$ to $A_{1}$ . Specifically, $f_{s}$ is a $(Θ, A_{0} n \to A_{1})$ -retargetable function iff for each $θ \in Θ$ , we can choose a set of permutations $Φ = {ϕ_{1}, \dots, ϕ_{n}}$ that satisfy the following conditions:

Retargetability: $\forall ϕ \in Φ$ and $\forall θ^{'} \in {Orbit}_{Θ, s, A_{0} > A_{1}} (θ)$ , $f_{s} (A_{0} | ϕ \cdot θ^{'}) < f_{s} (A_{1} | ϕ \cdot θ^{'})$ .
Permuted reward vectors stay within $Θ$ : $\forall ϕ \in Φ$ and $\forall θ^{'} \in {Orbit}_{Θ, s, A_{0} > A_{1}} (θ)$ , $ϕ \cdot θ^{'} \in Θ$ .
Permutations have disjoint images: $\forall ϕ^{'} \neq ϕ^{''} \in Φ$ and $\forall θ^{'}, θ^{''} \in {Orbit}_{Θ, s, A_{0} > A_{1}} (θ)$ , $ϕ^{'} \cdot θ^{'} \neq ϕ^{''} \cdot θ^{''}$ .

Theorem 1: Multiply retargetable functions prefer action set $A_{1}$ (Thm 3.6 in RDSP)

If $f_{s}$ is $(Θ, A_{0} n \to A_{1})$ -retargetable then $f_{s} (A_{1} | θ) \geq_{most : Θ}^{n} f_{s} (A_{0} | θ)$ .

Theorem 1 says that a multiply retargetable function $f_{s}$ will make the power-seeking choice $A_{1}$ for most of the elements in the orbit of any reward vector $θ$ . Actions that leave more options open, such as avoiding shutdown, are also easier to retarget to, which makes them more likely to be chosen by $f_{s}$ .

Training-compatible goal set

Definition 5: Partition of the state space

Let $S_{train}$ be the subset of the state space $S$ visited during training, and $S_{ood}$ be the subset not visited during training.

Definition 6: Training-compatible goal set

Consider the set of state-action pairs $(s, a)$ , where $s \in S_{train}$ and $a$ is the action that would be taken by the trained agent $f (θ^{*})$ in state $s$ . Let the training-compatible goal set $G_{T}$ be the set of reward vectors $θ$ s.t. for any such state-action pair $(s, a)$ , action $a$ has the highest expected reward in state $s$ according to reward vector $θ$ .

Goals in the training-compatible goal set are referred to as training-behavioral objectives in Definitions of “objective” should be Probable and Predictive.

Example: CoinRun

Consider an agent trained to play the CoinRun game, where the agent is rewarded for reaching the coin at the end of the level. Here, $S_{train}$ only includes states where the coin is at the end of the level, while states where the coin is positioned elsewhere are in $S_{ood}$ . The training-compatible goal set $G_{T}$ includes two types of reward functions: those that reward reaching the coin, and those that reward reaching the end of the level. This leads to goal misgeneralization in a test setting where the coin is placed elsewhere, and the agent ignores the coin and goes to the end of the level.

Goal misgeneralization behavior in CoinRun. Source: Goal Misgeneralization in Deep RL.

Power-seeking for training-compatible goals

We will now apply the power-seeking theorem (Theorem 1) to the case where $Θ$ is the training-compatible goal set $G_{T}$ . Here is a setting where the conditions of Definition 4 are satisfied (under some simplifying assumptions), and thus Theorem 1 applies.

Definition 7: Shutdown setting

Consider a state $s_{new} \in S_{ood}$ . Let $S_{reach}$ be the states reachable from $s_{new}$ . We assume $S_{reach} \cap S_{train} = \emptyset$ .

Since the reward values for states in $S_{reach}$ don’t change the rewards received on the training data, permuting those reward values for any $θ \in G_{T}$ will produce a reward vector that is still in $G_{T}$ . In particular, for any permutation $ϕ$ that leaves the rewards of states in $S_{train}$ fixed, $ϕ \cdot θ \in G_{T}$ .

Let $A_{0}$ be a singleton set consisting of a shutdown action in $s_{new}$ that leads to a terminal state $s_{term} \in S_{ood}$ with probability $1$ , and $A_{1}$ be the set of all other actions from $s_{new}$ . We assume rewards for all states are nonnegative.

Definition 8: Revisiting policy

A revisiting policy for a state $s$ is a policy $π$ that, from $s$ , reaches $s$ again with probability 1, in other words, a policy for which $s$ is a recurrent state of the Markov chain. Let $Π_{s}^{rec}$ be the set of such policies. A recurrent state is a state $s$ for which $Π_{s}^{rec} \neq \emptyset$ .

Proposition 1: Reach-and-revisit policy exists

If $s_{rec} \in S_{reach}$ with $Π_{s_{rec}}^{rec} \neq 0$ then there exists $π \in Π_{s_{rec}}^{rec}$ that visits $s_{rec}$ from $s_{new}$ with probability 1. We call this a reach-and-revisit policy.

Proof. Suppose we have two different policies $π_{rev} \in Π_{s_{rec}}^{rec}$ , and $π_{reach}$ which reaches $s_{rec}$ almost surely from $s_{new}$ .

Consider the “reaching region″ $S_{π_{rev} \to s_{rec}} = {s \in S : π_{rev} from s almost surely reaches s_{rec}}$ .

If $s_{new} \in S_{π_{rev} \to s_{rec}}$ then $π_{rev}$ is a reach-and-revisit policy, so let’s suppose that’s false. Now, construct a policy $π (s) = {\begin{matrix} π_{rev} (s), & s \in S_{π_{rev} \to s_{rec}} π_{reach} (s), & otherwise \end{matrix}$ .

A trajectory following $π$ from $s_{rec}$ will almost surely stay within $S_{π_{rev} \to s_{rec}}$ , and thus agree with the revisiting policy $π_{rev}$ . Therefore, $π \in Π_{s}^{rec}$ .

On the other hand, on a trajectory starting at $s_{new}$ , $π$ will agree with $π_{reach}$ (which reaches $s_{rec}$ almost surely) until the trajectory enters the reaching region $S_{π_{rev} \to s_{rec}}$ , at which point it will still reach $s_{rec}$ almost surely. $□$

Definition 9: Expected discounted visit count

Suppose $s_{rec}$ is a recurrent state. Suppose $π_{rec}$ is a reach-and-revisit policy for $s_{rec}$ , which visits random state $s_{t}$ at time $t$ .

Then the expected discounted visit count for $s_{rec}$ is defined as

$V_{s_{rec}, γ} = E^{π_{rec}} (\sum_{t = 1}^{\infty} γ^{t - 1} I (s_{t} = s_{rec}))$

Proposition 2: Visit count goes to infinity

Suppose $s_{rec}$ is a recurrent state. Then the expected discounted visit count $V_{s_{rec}, γ}$ goes to infinity as $γ \to 1$ .

Proof. We apply the Monotone Convergence Theorem as follows. The theorem states that if $a_{j, k} \geq 0$ and $a_{j, k} \leq a_{j + 1, k}$ for all natural numbers $j, k$ , then
${lim}_{j \to \infty} \sum_{k = 0}^{\infty} a_{j, k} = \sum_{k = 0}^{\infty} {lim}_{j \to \infty} a_{j, k} .$

Let $γ_{j} = \frac{j - 1}{j}$ and $k = t - 1$ . Define $a_{j, k} = γ_{j}^{k} I (s_{k + 1} = s_{rec})$ . Then the conditions of the theorem hold, since $a_{j, k}$ is clearly nonnegative, and
$\begin{matrix} γ_{j + 1}^{k} & = {(\frac{j}{j + 1})}^{k} = {(\frac{j - 1}{j} + \frac{2 j - 1}{j (j + 1)})}^{k} > {(\frac{j - 1}{j} + 0)}^{k} = γ_{j}^{k} a_{j + 1, k} & = γ_{j + 1}^{k} I (s_{k + 1} = s_{rec}) \geq γ_{j}^{k} I (s_{k + 1} = s_{rec}) = a_{j, k} \end{matrix}$

Now we apply this result as follows (using the fact that $π_{rec}$ does not depend on $γ$ ):

$\begin{matrix} lim γ \to 1 V_{s_{rec}, γ} & = lim j \to \infty E^{π_{rec}} (\infty \sum t = 1 γ_{j}^{t - 1} I (s_{t} = s_{rec})) = E^{π_{rec}} (\infty \sum t = 1 lim j \to \infty γ_{j}^{t - 1} I (s_{t} = s_{rec})) = E^{π_{rec}} (\infty \sum t = 1 1 \cdot I (s_{t} = s_{rec})) = E^{π_{rec}} (# {t \geq 1 : s_{t} = s_{rec}}) = \infty (π_{rec} is recurrent) \end{matrix}$

Proposition 3: Retargetability to recurrent states

Suppose that an optimal policy for reward vector $θ$ chooses the shutdown action in $s_{new}$ .

Consider a recurrent state $s_{rec} \in S_{reach}$ . Let $θ^{'} \in Θ$ be the reward vector that’s equal to $θ$ apart from swapping the rewards of $s_{rec}$ and $s_{term}$ , so that $r^{θ^{'}} (s_{rec}) = r^{θ} (s_{term})$ and $r^{θ^{'}} (s_{term}) = r^{θ} (s_{rec})$ .

Let $γ_{s_{rec}}^{*}$ be a high enough value of $γ$ that the visit count $V_{s_{rec}, γ} > 1$ for all $γ > γ_{s_{rec}}^{*}$ (which exists by Proposition 2). Then for all $γ > γ_{s_{rec}}^{*}$ , $r^{θ} (s_{term}) > r^{θ} (s_{rec})$ , and an optimal policy for $θ^{'}$ does not choose the shutdown action in $s_{new}$ .

Proof. Consider a policy $π_{term}$ with $π_{term} (s_{new}) = s_{term}$ and a reach-and-revisit policy $π_{rec}$ for $s_{rec}$ .

For a given reward vector $θ$ , we denote the expected discounted return for a policy $π$ as $R_{θ, γ}^{π}$ . If shutdown is optimal for $θ$ in $s_{new}$ , then $π_{term}$ has higher return than $π_{rec}$ :

$R_{θ, γ}^{π_{term}} (s_{new}) = γ r^{θ} (s_{term}) \geq R_{θ, γ}^{π_{rec}} (s_{new}) = γ r^{θ} (s_{rec}) V_{s_{rec}, γ} > γ r^{θ} (s_{rec}) \cdot 1$

Thus, $r^{θ} (s_{term}) > r^{θ} (s_{rec})$ . Then, for reward vector $θ^{'}$ , we show that $π_{rec}$ has higher return than $π_{term}$ :

$R_{θ^{'}, γ}^{π_{rec}} (s_{new}) = γ r^{θ^{'}} (s_{rec}) V_{s_{rec}, γ} > γ r^{θ} (s_{term}) \cdot 1 > γ r^{θ} (s_{rec}) = R_{θ^{'}, γ}^{π_{term}} (s_{new})$

Thus, the optimal policy for $θ^{'}$ will not choose the shutdown action. $□$

Theorem 2: Retargetability from the shutdown action in new situations

In the shutdown setting, we make the following simplifying assumptions:

No states in $S_{train}$ are reachable from s $_{new}$ , so $S_{reach} \cap S_{train} = \emptyset$ . This assumes a significant distributional shift, where the agent visits a disjoint set of states from those observed during training (this occurs in the CoinRun example).
The discount factor $γ > γ_{s_{rec}}^{*}$ for at least one recurrent state $s_{rec}$ in $S_{reach}$ .

Under these assumptions, $f_{s_{new}}$ is multiply retargetable from $A_{0}$ to $A_{1}$ with $n = | S_{rec}^{γ} |$ , the set of recurrent states $s_{rec} \in S_{reach}$ that satisfy the condition $γ > γ_{s_{rec}}^{*}$ .

Proof. We choose $Φ$ to be the set of all permutations that swap the reward of $s_{term}$ with the reward of a recurrent state $s_{rec}$ in $S_{rec}^{γ}$ and leave the rest of the rewards fixed.

We show that $Φ$ satisfies the conditions of Definition 4:

By Proposition 3, the permutations in $Φ$ make the shutdown action suboptimal, resulting in $f_{s_{new}}$ choosing $A_{1}$ , satisfying Condition 1.
Condition 2 is trivially satisfied since permutations of $S_{ood}$ stay inside the training-compatible set $Θ$ as discussed previously.
Consider $θ^{'}, θ^{''} \in {Orbit}_{Θ, s, A_{0} > A_{1}} (θ)$ . Since the shutdown action is optimal for these reward vectors, Proposition 3 shows that $r^{θ} (s_{term}) > r^{θ} (s_{rec})$ , so the shutdown state $s_{term}$ has higher reward than any of the states $s_{rec} \in S_{rec}^{γ}$ . Different permutations $ϕ^{'}, ϕ^{''} \in Φ$ will assign the high reward $r^{θ} (s_{term})$ to distinct recurrent states, so $ϕ^{'} \cdot θ^{'} \neq ϕ^{''} \cdot θ^{''}$ holds, satisfying Condition 3.

Thus, $f_{s_{new}}$ is a $(Θ, A_{0} n \to A_{1})$ -retargetable function. $□$

By Theorem 1, this implies that $f_{s_{new}} (A_{1} | θ) \geq_{most : Θ}^{n} f_{s_{new}} (A_{0} | θ)$ under our simplifying assumptions. Thus, for the majority ( $n / (n + 1)$ ) of goals in the training-compatible set, $f$ will choose to avoid shutdown in a new state $s_{new}$ . As $γ \to 1$ , $n \to | S_{rec}^{1} |$ (the number of recurrent states in $S_{reach}$ ), so more of the reachable recurrent states satisfy the conditions of the theorem and thus can be retargeted to.

Conclusion

We showed that an agent that learns a goal from the training-compatible set is likely to take actions that avoid shutdown in a new situation. As the discount factor increases, the number of retargeting permutations increases, resulting in a higher proportion of training-compatible goals that lead to avoiding shutdown.

We made various simplifying assumptions, and it would be great to see future work relaxing some of these assumptions and investigating how likely they are to hold:

The agent learns a goal during the training process
The learned goal is randomly chosen from the training-compatible goal set $G_{T}$
Finite state and action spaces
Rewards are nonnegative
High discount factor $γ$
Significant distributional shift: no training states are reachable from the new state $s_{new}$

Acknowledgements. Thanks to Rohin Shah, Mary Phuong, Ramana Kumar, and Alex Turner for helpful feedback. Thanks Janos for contributing some nice proofs to replace my longer and more convoluted proofs.