I’m talking about an imitation version where the human you’re imitating is allowed to do anything they want, including instatiting a search over all possible outputs X and taking that one that maximizes the score of “How good is answer X to Y?” to try to find X*. So I’m more pointing out that this behaviour is available in imitation by default. We could try to rule it out by instructing the human to only do limited searches, but that might be hard to do along with maintaining capabilities of the system, and we need to figure out what “safe limited search” actually looks like.
I’m talking about an imitation version where the human you’re imitating is allowed to do anything they want, including instatiting a search over all possible outputs X and taking that one that maximizes the score of “How good is answer X to Y?” to try to find X*. So I’m more pointing out that this behaviour is available in imitation by default. We could try to rule it out by instructing the human to only do limited searches, but that might be hard to do along with maintaining capabilities of the system, and we need to figure out what “safe limited search” actually looks like.