Jozdien comments on Run evals on base models too!

Jozdien 4 Apr 2024 21:09 UTC
2 points
0
I agree, and one of the reasons I endorse this strongly: there’s a difference between performance at a task in a completion format and a chat format^[1].
Chess, for instance, is most commonly represented in the training data in PGN format—and prompting it with something that’s similar should elicit whatever capabilities it possesses more strongly than transferring them over to another format.
I think this should generalize pretty well across more capabilities—if you expect the bulk of the capability to come from pre-training, then I think you should also expect these capabilities to be most salient in completion formats.
1. ^
  Everything in a language model is technically in completion format all the time, but the distinction I want to draw here is related to tasks that are closest to the training data (and therefore closest to what powerful models are superhuman at). Models are superhuman at next-token prediction, and inasmuch as many examples of capabilities in the training data are given in prose as opposed to a chat format (for example, chess), you should expect those capabilities to be most salient in the former.