Skip to content
SI.info

SI Glossary · Training & data

Synthetic Data

Published 1 min read
On this page
  1. Why labs use it
  2. Risks
  3. Related: distillation

Synthetic data is artificially generated: questions and answers written by a model, simulated driving scenes, or computer-generated images with perfect labels. It has become a major ingredient in training frontier SI.

Why labs use it

  • Scarcity: good human-written examples of advanced maths, coding or reasoning are limited.
  • Privacy: synthetic patient or customer records can stand in for sensitive real data.
  • Control: developers can generate exactly the cases they need, including rare edge cases.
  • Verification: in maths and code, synthetic problems can be checked automatically, which makes them ideal for training reasoning models.

Risks

Training models mostly on other models’ output can amplify errors and narrow diversity, an effect researchers call “model collapse” when it compounds over generations. Labs mix synthetic data with real data and filter it carefully.

When a smaller model is trained on a larger model’s outputs, that is a form of synthetic data called distillation.

← Back to the SI Glossary

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy