But creating extreme suffering might not actually involve doing much in the physical world (compared to “normal” actions the AI would have to take to achieve the goals that we gave it). What if, depending on the goals we give the AI, doing this kind of extortion is actually the lowest impact way to achieve some goal?
Since there are a lot of possible scenarios, each of which affects the optimization differently, I’m hesitant to use a universal quantifier here without more details. However, I am broadly suspicious of AUP agents choosing plans which involve almost maximally offensive components, even accounting for the fact that it could try to do so surreptitiously. An agent might try to extort us if it expected we would respond, but respond with what? Although impact measures quantify things in the environment, that doesn’t mean they’re measuring how “similar” two states look to the eye. AUP penalizes distance traveled in the Q function space for its attainable utility functions. We also need to think about the motive for the extortion – if it means the agent gains in power, then that is also penalized.
Maybe it could extort a different group of humans, and as part of the extortion force them to keep it secret from people who could turn it off? Or extort us and as part of the extortion force us to not turn it off (until we were going to turn it off anyway)?
Again, it depends on the objective of the extortion. As for the latter, that wouldn’t be credible, since we would be able to tell its threat was the last action in its plan. AUP isolates the long-term effects of each action by having the agent stop acting for the rest of the epoch; this gives us a counterfactual opportunity to respond to that action.
I’m not sure whether this belongs in the desiderata, since we’re talking about whether temporary object level bad things could happen. I think it’s a bonus to think that there is less of a chance of that, but not the primary focus of the impact measure. Even so, it’s true that we could explicitly talk about what we want to do with impact measures, adding desiderata like “able to do reasonable things” and “disallows catastrophes from rising to the top of the preference ordering”. I’m still thinking about this.
However, I am broadly suspicious of AUP agents choosing plans which involve almost maximally offensive components, even accounting for the fact that it could try to do so surreptitiously.
I guess I don’t have good intuitions of what an AUP agent would or wouldn’t do. Can you share yours, like give some examples of real goals we might want to give to AUP agents, and what you think they would and wouldn’t do to accomplish each of those goals, and why? (Maybe this could be written up as a post since it might be helpful for others to understand your intuitions about how AUP would work in a real-world setting.)
I’m not sure whether this belongs in the desiderata, since we’re talking about whether temporary object level bad things could happen. I think it’s a bonus to think that there is less of a chance of that, but not the primary focus of the impact measure.
Why not? I’ve usually seen people talk about “impact measures” as a way of avoiding side effects, especially negative side effects. It seems intuitive that “object level bad things” are negative side effects even if they are temporary, and ought to be a primary focus of impact measures. It seems like you’ve reframed “impact measures” in your mind to be a bit different from this naive intuitive picture, so perhaps you could explain that a bit more (or point me to such an explanation)?
Since there are a lot of possible scenarios, each of which affects the optimization differently, I’m hesitant to use a universal quantifier here without more details. However, I am broadly suspicious of AUP agents choosing plans which involve almost maximally offensive components, even accounting for the fact that it could try to do so surreptitiously. An agent might try to extort us if it expected we would respond, but respond with what? Although impact measures quantify things in the environment, that doesn’t mean they’re measuring how “similar” two states look to the eye. AUP penalizes distance traveled in the Q function space for its attainable utility functions. We also need to think about the motive for the extortion – if it means the agent gains in power, then that is also penalized.
Again, it depends on the objective of the extortion. As for the latter, that wouldn’t be credible, since we would be able to tell its threat was the last action in its plan. AUP isolates the long-term effects of each action by having the agent stop acting for the rest of the epoch; this gives us a counterfactual opportunity to respond to that action.
I’m not sure whether this belongs in the desiderata, since we’re talking about whether temporary object level bad things could happen. I think it’s a bonus to think that there is less of a chance of that, but not the primary focus of the impact measure. Even so, it’s true that we could explicitly talk about what we want to do with impact measures, adding desiderata like “able to do reasonable things” and “disallows catastrophes from rising to the top of the preference ordering”. I’m still thinking about this.
I guess I don’t have good intuitions of what an AUP agent would or wouldn’t do. Can you share yours, like give some examples of real goals we might want to give to AUP agents, and what you think they would and wouldn’t do to accomplish each of those goals, and why? (Maybe this could be written up as a post since it might be helpful for others to understand your intuitions about how AUP would work in a real-world setting.)
Why not? I’ve usually seen people talk about “impact measures” as a way of avoiding side effects, especially negative side effects. It seems intuitive that “object level bad things” are negative side effects even if they are temporary, and ought to be a primary focus of impact measures. It seems like you’ve reframed “impact measures” in your mind to be a bit different from this naive intuitive picture, so perhaps you could explain that a bit more (or point me to such an explanation)?
Sounds good. I’m currently working on a long sequence walking through my intuitions and assumptions in detail.