Could it be inefficient scaling? Most work not explicitly using scaling laws to plan it seems to generally overestimate in compute per parameter, using too-small models. Anyone want to try to apply Jones 2021 to see if AlphaZero was scaled wrong?
Ben Adlam (via Maloney et al 2022) makes an interesting point: if you plot parameters vs training data, it’s a nearly perfect 1:1 ratio historically. (He doesn’t seem to have published anything formally on this.)
We have conveniently just updated our database if anyone wants to investigate this further!https://epochai.org/data/notable-ai-models
Could it be inefficient scaling? Most work not explicitly using scaling laws to plan it seems to generally overestimate in compute per parameter, using too-small models. Anyone want to try to apply Jones 2021 to see if AlphaZero was scaled wrong?
Ben Adlam (via Maloney et al 2022) makes an interesting point: if you plot parameters vs training data, it’s a nearly perfect 1:1 ratio historically. (He doesn’t seem to have published anything formally on this.)
We have conveniently just updated our database if anyone wants to investigate this further!
https://epochai.org/data/notable-ai-models