sudo comments on How much should we worry about mesa-optimization challenges?

sudo 28 Jul 2022 6:16 UTC
2 points
−1
Sorry, I’ve been too busy to reply. I’m still too busy to give an incredibly detailed reply, but I can at least give a reply. A reply is better than no reply.
An unstable goal leads to near-term behavior that fits different utility functions as it changes.
“It changes to fit different utility functions” is not distinguishable from “it has a single, complex, persistent utility function which rewards drastically differing policies in incredibly similar but subtly different contexts.” An agent is never in the exact same environment twice.
So an agent with unstable goals is not an optimizer for any goal, other than indirectly for its CEV where it path-independently tends to eventually go, but its current behavior doesn’t hint at that yet, doesn’t even mildly optimize for CEV taken as a goal, because it doesn’t yet know what its CEV is.
This framing seems significant and important to you. However, I fail to see its utility. Could you help me see why this is how you chose to look at the problem?
- Vladimir_Nesov 28 Jul 2022 6:53 UTC
  4 points
  3
  Parent
  What serves as a goal in distant future determines how cosmic endowment is optimized. Stable goals are also goals that remain in distant future, so they are relevant to that (and since reflection hasn’t yet had a chance of having taken place, stable goals settled in near future are always misaligned). Unstable goals are not relevent in themselves, in what utility function (or maybe probutility) they fit, except in how they tend to produce different stable goals eventually.
  
  So maintaining the distinction means not being unaware of the catastrophic misalignment risk where we turn some unstable goals into stable ones based on a stupid process of (possibly lack of) reflection that just fits things instead of doing proper well-designed reflection (a thing like CEV, possibly very different in detail). And it helps with not worrying too much about details of utility functions that fit current unstable goals, or aligning them with human current unstable goals, when they are not what actually matters.
  
  An agent is never in the exact same environment twice.
  
  That doesn’t affect goals, which talk of all possible environments, doesn’t matter if some agent actually encounters them. Goals are not just policy, instead they determine policy, not the other way around (along the algorithm vs. physical distinction, goals are closer to the algorithm, while policy is merely the behavior of the algorithm, the decision taken by it, closer to the physical instances and actions in reality). Unstable goals change their mind about the same environment. It could be an environment that will be reachable/enactable in the future.