Vladimir_Nesov comments on O O’s Shortform

Vladimir_Nesov 17 Nov 2024 19:39 UTC
2 points
0

for anything related to human judgement, in theory this isn’t why it’s not doing well

The facts are in there, but not in the form of a sufficiently good reward model that can tell as well as human experts which answer is better or whether a step of an argument is valid. In the same way, RLHF is still better with humans on some queries, hasn’t been fully automated to superior results by replacing humans with models in all cases.