More than a year since writing this post, I would still say it represents the key ideas in the sequence on mesa-optimisation which remain central in today’s conversations on mesa-optimisation. I still largely stand by what I wrote, and recommend this post as a complement to that sequence for two reasons:
First, skipping some detail allows it to focus on the important points, making it better-suited than the full sequence for obtaining an overview of the area.
Second, unlike the sequence, it deemphasises the mechanism of optimisation, and explicitly casts it as a way of talking about goal-directedness. As time passes, I become more and more convinced that it was a mistake to call the primary new term in our work ‘mesa-optimisation’. Were I to be choosing the terms again, I would probably go with something like ‘learned goal-directedness’, though it is quite a mouthful.
More than a year since writing this post, I would still say it represents the key ideas in the sequence on mesa-optimisation which remain central in today’s conversations on mesa-optimisation. I still largely stand by what I wrote, and recommend this post as a complement to that sequence for two reasons:
First, skipping some detail allows it to focus on the important points, making it better-suited than the full sequence for obtaining an overview of the area.
Second, unlike the sequence, it deemphasises the mechanism of optimisation, and explicitly casts it as a way of talking about goal-directedness. As time passes, I become more and more convinced that it was a mistake to call the primary new term in our work ‘mesa-optimisation’. Were I to be choosing the terms again, I would probably go with something like ‘learned goal-directedness’, though it is quite a mouthful.