Skip to content
SI.info

SI Glossary · Training & data

Pretraining

Published 1 min read
On this page
  1. What happens during pretraining
  2. Cost and scale
  3. What comes after

Pretraining is where a foundation model gets most of its knowledge. For a large language model, it means reading trillions of tokens of text and code and learning to predict the next token, over and over, for weeks or months on thousands of chips.

What happens during pretraining

  • The model starts with random parameters.
  • It sees a chunk of text, predicts the next token, and is corrected.
  • Repeated trillions of times, this forces it to learn grammar, facts, reasoning patterns, coding conventions and much more.

This is called self-supervised learning: no human labels are needed, because the text supplies its own answers.

Cost and scale

Frontier pretraining runs use enormous amounts of compute and cost hundreds of millions of dollars or more, mostly in GPUs and electricity. That expense is why laws often use training compute as a trigger for safety rules. See compute threshold.

What comes after

A freshly pretrained model is knowledgeable but not yet a helpful assistant. It might continue your question with more questions. Post-training (fine-tuning, RLHF and reinforcement learning for reasoning) turns it into a usable product. Since about 2024, labs have shifted more of their budgets toward post-training and reasoning, while continuing to scale pretraining.

← Back to the SI Glossary

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy