Skip to content
SI.info

SI Glossary · Training & data

Scaling Laws

Published 1 min read
On this page
  1. Key findings
  2. Beyond pretraining
  3. Will scaling continue?

Scaling laws describe one of the most important discoveries in modern SI: model performance improves smoothly and predictably as you add more compute, more training data and more parameters.

Key findings

  • 2020, OpenAI (Kaplan et al.): language-model loss falls as a power law with model size, dataset size and compute, across many orders of magnitude.
  • 2022, DeepMind (“Chinchilla”): for a fixed compute budget, many models were too big and under-trained. Roughly, parameters and training tokens should grow together.

These results gave labs confidence to spend ever larger sums on training runs, and help explain the boom in GPU and data center investment.

Beyond pretraining

Since 2024, labs have found new scaling axes:

Will scaling continue?

This is the central bet of the industry. Optimists say scaling plus new techniques leads to AGI and beyond. Skeptics point to rising costs, data limits and the gap between benchmark gains and real-world reliability. See When will superintelligence arrive?

← Back to the SI Glossary

Sources

  1. Scaling Laws for Neural Language Models — Kaplan et al., OpenAI, 2020
  2. Training Compute-Optimal Large Language Models (Chinchilla) — Hoffmann et al., DeepMind, 2022

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy