SI Glossary · Training & data
Distillation
Distillation compresses the abilities of a big “teacher” model into a smaller “student”. The student is trained on the teacher’s outputs (its answers, reasoning or probability scores) rather than only on raw data. The idea was popularised by Geoffrey Hinton and colleagues in 2015.
Why it matters
- Cheaper, faster models: many “mini”, “flash” and “lite” models are distilled from flagship models.
- On-device SI: small distilled models can run on phones and laptops.
- Open weights: labs release distilled versions of large open models that more people can run. Thinking Machines’ Inkling Small (2026) is described as a distilled version of Inkling.
Controversy
Distilling someone else’s model through its API usually violates the provider’s terms of service. Accusations that rivals distilled frontier models have become a recurring flashpoint between U.S. and Chinese labs. Policymakers have also debated whether distilled models should count as frontier models: New York’s RAISE Act originally covered them, but the March 2026 amendment dropped that.
Sources
- Distilling the Knowledge in a Neural Network — Hinton, Vinyals & Dean, 2015
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.