At first, the AI would converge towards: “my reward button corresponds to (is) doing what humans want”, and that conceptualization would become the centerpiece, so to speak, of its reasoning ability: the locus through which everything is filtered. The thought of pressing the reward button directly, bypassing humans, would also be filtered into that initial reward-conception… which would reject it offhand. So even though the AI is getting smarter and smarter, it is hopelessly stuck in a local minimum and expends no energy getting out of it.
This is a Value Learner, not a Reinforcement Learner like the standard AIXI. They’re two different agent models, and yes, Value Learners have been considered as tools for obtaining an eventual Seed AI. I personally (ie: massive grains of salt should be taken by you) find it relatively plausible that we could use a Value Learner as a Tool AGI to help us build a Friendly Seed AI that could then be “unleashed” (ie: actually unboxed and allowed into the physical universe).
This is a Value Learner, not a Reinforcement Learner like the standard AIXI. They’re two different agent models, and yes, Value Learners have been considered as tools for obtaining an eventual Seed AI. I personally (ie: massive grains of salt should be taken by you) find it relatively plausible that we could use a Value Learner as a Tool AGI to help us build a Friendly Seed AI that could then be “unleashed” (ie: actually unboxed and allowed into the physical universe).