Skip to content
SI.info

SI Glossary · Models & architecture

Mixture of Experts (MoE)

Published 1 min read
On this page
  1. Why it matters
  2. Trade-offs
  3. History

A mixture-of-experts (MoE) model contains many “expert” sub-networks plus a small router that decides which experts to use for each token. Only a fraction of the model does work at any moment.

Why it matters

MoE separates total size from running cost:

  • A model can store a huge amount of knowledge in its total parameters.
  • Each token only pays for the active parameters, so inference is faster and cheaper.

For example, Alibaba’s open-weights Qwen3.8-Max is reported at about 2.4 trillion total parameters with roughly 95 billion active per token, and DeepSeek’s V4 family also uses MoE. See our model tracker.

Trade-offs

MoE models need a lot of memory, because all experts must be loaded even if few are used. Training is harder to balance, because routers can overuse some experts. That’s why “you can download it” doesn’t always mean “you can run it” for the largest open-weights models.

History

The idea dates to 1991 research by Robert Jacobs, Michael Jordan, Steven Nowlan and Geoffrey Hinton. Google’s 2017 “sparsely-gated mixture-of-experts” paper brought it to large neural networks, and it became mainstream in frontier language models in the mid-2020s.

← Back to the SI Glossary

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy