They’re also finding that inverse scaling on these tasks goes away with chain-of-thought prompting
So, like some of the Big-Bench PaLM results, these are more cases of ‘hidden scaling’ where quite simple inner-monologue approaches can show smooth scaling while the naive pre-existing benchmark claims that there are no gains with scale?
So, like some of the Big-Bench PaLM results, these are more cases of ‘hidden scaling’ where quite simple inner-monologue approaches can show smooth scaling while the naive pre-existing benchmark claims that there are no gains with scale?
Yup