TurnTrout comments on TurnTrout’s shortform feed

TurnTrout 18 Mar 2020 18:36 UTC
LW: 4 AF: 2
AF
Very rough idea

In 2018, I started thinking about corrigibility as “being the kind of agent lots of agents would be happy to have activated”. This seems really close to a more ambitious version of what AUP tries to do (not be catastrophic for most agents).

I wonder if you could build an agent that rewrites itself / makes an agent which would tailor the AU landscape towards its creators’ interests, under a wide distribution of creator agent goals/rationalities/capabilities. And maybe you then get a kind of generalization, where most simple algorithms which solve this solve ambitious AI alignment in full generality.