SI Glossary · Training & data
Training Data
On this page
Training data is everything a model learns from. For frontier large language models, that means trillions of tokens drawn from:
- the public web (filtered heavily for quality),
- books, academic papers and code repositories,
- licensed content from publishers and data providers,
- human-written examples and ratings for fine-tuning and RLHF,
- increasingly, synthetic data generated by other models.
Why it matters
- Quality beats quantity: careful filtering and deduplication improve models as much as raw size does.
- Knowledge cutoff: a model knows little about events after its training data ends.
- Bias: models absorb the patterns, including stereotypes, present in their data.
The legal fight
Whether training on copyrighted material without permission is lawful has been one of the biggest legal questions of the era, with many lawsuits filed by authors, artists, news organisations and music companies. The U.S. administration’s March 2026 National Policy Framework treats training on copyrighted material as lawful while encouraging voluntary licensing systems for creators. Courts and other countries may see it differently.
Running out of data?
Researchers have warned that high-quality human text is finite. Labs are responding with synthetic data, licensing deals, multimodal data (video, audio) and reinforcement learning that relies less on new text.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.