Wouldn’t really need reward modelling for narrow optimizers. Weak general real-world optimizers, I find difficult to imagine, and I’d expect them to be continuous with strong ones, the projects to make weak ones wouldn’t be easily distinguishable from the projects to make strong ones.
Oh, are you thinking of applying it to say, simulation training.
Wouldn’t really need reward modelling for narrow optimizers. Weak general real-world optimizers, I find difficult to imagine, and I’d expect them to be continuous with strong ones, the projects to make weak ones wouldn’t be easily distinguishable from the projects to make strong ones.
Oh, are you thinking of applying it to say, simulation training.