Scaling laws and Chinchilla
Model quality moves predictably with compute, and the field spent years splitting that compute the wrong way.
Part of the How models are made track on lAItest.
For a while the field competed on parameter count. Then a 70B model beat a 530B one at equal training compute.
Scaling laws: the loss curve is smooth and it is predictable.
Plot how well a model predicts text against the compute, the data and the parameters used to train it and you do not get noise. You get smooth curves that hold across orders of magnitude. That is what turns a nine-figure training run into a planning exercise rather than a gamble: fit the curve on small runs, forecast where the big one lands. The reference work is arXiv 2001.08361.
Chinchilla: the split between size and data was wrong.
A fixed compute budget forces a choice: a bigger model on less data, or a smaller model on more. The Chinchilla work, arXiv 2203.15556, showed the field had been buying parameters and starving models of data. A 70B model trained on roughly four times more data outperformed Gopher, GPT-3 at 175B, Jurassic-1 at 178B and MT-NLG at 530B at equal compute, reaching 67.5% on MMLU. The rule of thumb people took away was around twenty training tokens per parameter.
A common misconception
Commonly believed: So the optimal recipe is twenty tokens per parameter, and anything else wastes compute.
Actually: Chinchilla optimises training compute alone. It says nothing about what the model costs to run afterwards. If a model will serve requests for years then training is a one-off and inference dominates lifetime cost, so labs deliberately overtrain, pushing far past twenty tokens per parameter to get a smaller model of the same quality. That is not a violation of the result. It is optimising a different objective.
Why do labs now train smaller models on far more data than the Chinchilla ratio suggests?
Answer: A smaller model is cheaper to run, and inference dominates lifetime cost. Chinchilla answers one question: what is the best model I can train for this compute budget. Deployment asks a different one: what is the cheapest model that reaches this quality and then serves traffic for two years. Overtraining a smaller model spends more up front to lower every request bill afterwards.
In one sentence
Chinchilla told the field it was building models too big and feeding them too little, and modern practice overshoots the correction on purpose.