SI Glossary · Training & data
Scaling Laws
On this page
Scaling laws describe one of the most important discoveries in modern SI: model performance improves smoothly and predictably as you add more compute, more training data and more parameters.
Key findings
- 2020, OpenAI (Kaplan et al.): language-model loss falls as a power law with model size, dataset size and compute, across many orders of magnitude.
- 2022, DeepMind (“Chinchilla”): for a fixed compute budget, many models were too big and under-trained. Roughly, parameters and training tokens should grow together.
These results gave labs confidence to spend ever larger sums on training runs, and help explain the boom in GPU and data center investment.
Beyond pretraining
Since 2024, labs have found new scaling axes:
- Post-training compute: reinforcement learning for reasoning.
- Test-time compute: letting a model think longer at inference produces better answers.
Will scaling continue?
This is the central bet of the industry. Optimists say scaling plus new techniques leads to AGI and beyond. Skeptics point to rising costs, data limits and the gap between benchmark gains and real-world reliability. See When will superintelligence arrive?
Sources
- Scaling Laws for Neural Language Models — Kaplan et al., OpenAI, 2020
- Training Compute-Optimal Large Language Models (Chinchilla) — Hoffmann et al., DeepMind, 2022
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.