SI Glossary · Training & data
Pretraining
Pretraining is where a foundation model gets most of its knowledge. For a large language model, it means reading trillions of tokens of text and code and learning to predict the next token, over and over, for weeks or months on thousands of chips.
What happens during pretraining
- The model starts with random parameters.
- It sees a chunk of text, predicts the next token, and is corrected.
- Repeated trillions of times, this forces it to learn grammar, facts, reasoning patterns, coding conventions and much more.
This is called self-supervised learning: no human labels are needed, because the text supplies its own answers.
Cost and scale
Frontier pretraining runs use enormous amounts of compute and cost hundreds of millions of dollars or more, mostly in GPUs and electricity. That expense is why laws often use training compute as a trigger for safety rules. See compute threshold.
What comes after
A freshly pretrained model is knowledgeable but not yet a helpful assistant. It might continue your question with more questions. Post-training (fine-tuning, RLHF and reinforcement learning for reasoning) turns it into a usable product. Since about 2024, labs have shifted more of their budgets toward post-training and reasoning, while continuing to scale pretraining.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.