Skip to content
SI.info

SI Glossary · Training & data

Training Data

Published 1 min read
On this page
  1. Why it matters
  2. The legal fight
  3. Running out of data?

Training data is everything a model learns from. For frontier large language models, that means trillions of tokens drawn from:

  • the public web (filtered heavily for quality),
  • books, academic papers and code repositories,
  • licensed content from publishers and data providers,
  • human-written examples and ratings for fine-tuning and RLHF,
  • increasingly, synthetic data generated by other models.

Why it matters

  • Quality beats quantity: careful filtering and deduplication improve models as much as raw size does.
  • Knowledge cutoff: a model knows little about events after its training data ends.
  • Bias: models absorb the patterns, including stereotypes, present in their data.

Whether training on copyrighted material without permission is lawful has been one of the biggest legal questions of the era, with many lawsuits filed by authors, artists, news organisations and music companies. The U.S. administration’s March 2026 National Policy Framework treats training on copyrighted material as lawful while encouraging voluntary licensing systems for creators. Courts and other countries may see it differently.

Running out of data?

Researchers have warned that high-quality human text is finite. Labs are responding with synthetic data, licensing deals, multimodal data (video, audio) and reinforcement learning that relies less on new text.

← Back to the SI Glossary

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy