sej2020

Karma: 59

sej2020 Dec 30, 2024, 7:11 PM
6 points
0
in reply to: FeepingCreature’s comment on: Greedy-Advantage-Aware RLHF
My thinking is not very clear on this point, but I am generally pessimistic that any type of RL/optimization regime with an adversarial nature could be robust to self-aware agents. To me, it seems like adversarial methodologies could spawn opposing mesaoptimizers, and we would be at the mercy of whichever subsystem represented its optimization process well enough to squash the other.

sej2020 Dec 30, 2024, 7:01 PM
4 points
0
in reply to: Charlie Steiner’s comment on: Greedy-Advantage-Aware RLHF
Martin is representing my claim well in this exchange, but I also think it’s important to mention that the simple/convoluted plan continuum does not have a perfect correspondence with the sharp/flat policy continuum. For example, wireheading may be simple in abstract, but I still expect a wireheading policy to be extremely sharp. If a wireheading policy takes, let’s say 5 distinct actions (WWWWW) to execute, and the agent’s policy is WWWWW, then it would receive arbitrarily high reward because the agent controls the reward button. However, if the policy is similar, like WWWWX, it would receive much lower reward because a plan that is 80% wireheading would likely not score well on the reward function representing the true goal.

sej2020 Dec 30, 2024, 6:49 PM
2 points
0
in reply to: Sam Marks’s comment on: Greedy-Advantage-Aware RLHF
Thanks for the feedback, these suggestions are definitely helpful as I’m thinking about how/if to advance the project.