SI Glossary · Training & data
Synthetic Data
On this page
Synthetic data is artificially generated: questions and answers written by a model, simulated driving scenes, or computer-generated images with perfect labels. It has become a major ingredient in training frontier SI.
Why labs use it
- Scarcity: good human-written examples of advanced maths, coding or reasoning are limited.
- Privacy: synthetic patient or customer records can stand in for sensitive real data.
- Control: developers can generate exactly the cases they need, including rare edge cases.
- Verification: in maths and code, synthetic problems can be checked automatically, which makes them ideal for training reasoning models.
Risks
Training models mostly on other models’ output can amplify errors and narrow diversity, an effect researchers call “model collapse” when it compounds over generations. Labs mix synthetic data with real data and filter it carefully.
Related: distillation
When a smaller model is trained on a larger model’s outputs, that is a form of synthetic data called distillation.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.