they almost certainly don’t have anything to do with what humans want, per se. (that would be basically magic)
We are obviously not appealing to literal telepathy or magic. Deep learning generalizes the way we want in part because we designed the architectures to be good, in part because human brains are built on similar principles to deep learning, and in part because we share a world with our deep learning models and are exposed to similar data.
I just deny that they will update “arbitrarily” far from the prior, and I don’t know why you would think otherwise. There are compute tradeoffs and you’re doing to run only as many MCTS rollouts as you need to get good performance.