There’s no guarantee that such a thing even exists, and implicitly aiming to avoid the one value formation process we know is compatible with our own values seems like a terrible idea.
...
It’s thus vastly easier to align models to goals where we have many examples of people executing said goals.
I think there’s a deep disconnect here on whether interpolation is enough or whether we need extrapolation.
The point of the strawberry alignment problem is “here’s a clearly understandable specification of a task that requires novel science and engineering to execute on. Can you do that safely?”. If your ambitions are simply to have AI customer service bots, you don’t need to solve this problem. If your ambitions include cognitive megaprojects which will need to be staffed at high levels by AI systems, then you do need to solve this problem.
More pragmatically, if your ambitions include setting up some sort of system that prevents people from deploying rogue AI systems while not dramatically curtailing Earth’s potential, that isn’t a goal that we have many examples of people executing on. So either we need to figure it out with humans or, if that’s too hard, create an AI system capable of figuring it out (which probably requires an AI leader instead of an AI assistant).
I think there’s a deep disconnect here on whether interpolation is enough or whether we need extrapolation.
The point of the strawberry alignment problem is “here’s a clearly understandable specification of a task that requires novel science and engineering to execute on. Can you do that safely?”. If your ambitions are simply to have AI customer service bots, you don’t need to solve this problem. If your ambitions include cognitive megaprojects which will need to be staffed at high levels by AI systems, then you do need to solve this problem.
More pragmatically, if your ambitions include setting up some sort of system that prevents people from deploying rogue AI systems while not dramatically curtailing Earth’s potential, that isn’t a goal that we have many examples of people executing on. So either we need to figure it out with humans or, if that’s too hard, create an AI system capable of figuring it out (which probably requires an AI leader instead of an AI assistant).