paulfchristiano comments on Corrigibility thoughts II: the robot operator

paulfchristiano 28 Jun 2017 15:10 UTC
0 points
AF
The AI defers to anything that can control the operator.

If the operator has physical control over the AI, than any process which controls the operator can replace the AI wholesale. It feels fine to defer to such processes, and certainly it seems much better than the situation where the operator is attempting to correct the AI’s behavior but the AI is paternalistically unresponsive.

Presumably the operator will try to secure themselves in the same way that they try to secure their AI.
- Stuart_Armstrong 29 Jun 2017 8:52 UTC
  0 points
  AF Parent
  This also means that if the AI can figure out a way of controlling the controller, then it is itself in control form the moment it comes up with a reasonable plan?
  - paulfchristiano 29 Jun 2017 17:14 UTC
    0 points
    AF Parent
    The AI replacing the operator is certainly a fixed point.
    
    This doesn’t seem any different from the usual situation. Modifying your goals is always a fixed point. That doesn’t mean that our agents will inevitably do it.
    
    An agent which is doing what the operator wants, where the operator is “whatever currently has physical control of the AI,” won’t try to replace the operator—because that’s not what the operator wants.
    - Stuart_Armstrong 30 Jun 2017 13:03 UTC
      0 points
      AF Parent
      
      An agent which is doing what the operator wants, where the operator is “whatever currently has physical control of the AI,” won’t try to replace the operator—because that’s not what the operator wants.
      
      I disagree (though we may be interpreting that sentence differently). Once the AI has the possibility of subverting the controller, then it is, in effect, in physical control of itself. So it itself becomes the “formal operator”, and, depending on how it’s motivated, is perfectly willing to replace the “human operator”, whose wishes are now irrelevant (because it’s no longer the formal operator).
      
      And this never involves any goal modification at all—it’s the same goal, except that the change in control has changed the definition of the operator.