Karl comments on Introducing Corrigibility (an FAI research subfield)

Karl 24 Oct 2014 2:38 UTC
6 points
Why not make it so that the agent in selecting A1 act as a UN-agent that believe that it will continue to optimize according to UN even in the event of the button being pressed rather than a UN agent that believe that the button will never be pressed: that is pick U such that

U(a1,o,a2) = UN(a1,o,a2) if o is in Press or US(a1,o,a2) + f(a1,o) - g(a1,o) if o is not in Press

where f(a1,o) is the maximum value of UN(a1,o,b) for b in A2 and g(a1,o) is the maximum value of US(a1,o,b) for b in A2.

This would avoid the perverse manipulation incentives problem detailed on section 4.2 of the paper.
- interstice 24 Oct 2014 3:13 UTC
  2 points
  Parent
  How does this differ from indifference?
  - Karl 24 Oct 2014 4:00 UTC
    3 points
    Parent
    In the indifference formalism the agent in selecting A1 act like a UN agent that believe that the shutdown button will not be pressed, therefore it create perverse incentives to “manage the news”. Which means that if the agent can cause his shutdown button to be pressed in the event of bad news, it will.
    
    My formulation avoid this pathological behavior by instead making the agent select A1 as if it was a UN-agent which believed that it would continue to optimize according to UN even in the vent of the button being pressed which avoid the perverse incentives to “manage the news”, while still not having any incentives to avoid the button being pressed because the agent will act like it believe that pressing the button will not cause it to initiate a shutdown.