Donald Hobson has a good point about goodharting, but it’s simpler than that. While some people want alignment so that everyone doesn’t die, the rest of us still want it for what it can do for us. If I’m prompting a language model with “A small cat went out to explore the world” I want it to come back with a nice children’s story about a small cat that went out to explore the world that I can show to just about any child. If I prompt a robot that I want it to “bring me a nice flower” I do not want it to steal my neighbor’s rosebushes. And so on, I want it to be safe to give some random AI helper whatever lazy prompt is on my mind and have it improve things by my preferences.
Donald Hobson has a good point about goodharting, but it’s simpler than that. While some people want alignment so that everyone doesn’t die, the rest of us still want it for what it can do for us. If I’m prompting a language model with “A small cat went out to explore the world” I want it to come back with a nice children’s story about a small cat that went out to explore the world that I can show to just about any child. If I prompt a robot that I want it to “bring me a nice flower” I do not want it to steal my neighbor’s rosebushes. And so on, I want it to be safe to give some random AI helper whatever lazy prompt is on my mind and have it improve things by my preferences.