Skip to content
SI.info

SI Glossary · Training & data

Distillation

Published Updated 1 min read

Distillation compresses the abilities of a big “teacher” model into a smaller “student”. The student is trained on the teacher’s outputs (its answers, reasoning or probability scores) rather than only on raw data. The idea was popularised by Geoffrey Hinton and colleagues in 2015.

Why it matters

  • Cheaper, faster models: many “mini”, “flash” and “lite” models are distilled from flagship models.
  • On-device SI: small distilled models can run on phones and laptops.
  • Open weights: labs release distilled versions of large open models that more people can run. Thinking Machines’ Inkling Small (2026) is described as a distilled version of Inkling.

Controversy

Distilling someone else’s model through its API usually violates the provider’s terms of service. Accusations that rivals distilled frontier models have become a recurring flashpoint between U.S. and Chinese labs. Policymakers have also debated whether distilled models should count as frontier models: New York’s RAISE Act originally covered them, but the March 2026 amendment dropped that.

← Back to the SI Glossary

Sources

  1. Distilling the Knowledge in a Neural Network — Hinton, Vinyals & Dean, 2015

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy