SI Glossary · Models & architecture
Mixture of Experts (MoE)
On this page
A mixture-of-experts (MoE) model contains many “expert” sub-networks plus a small router that decides which experts to use for each token. Only a fraction of the model does work at any moment.
Why it matters
MoE separates total size from running cost:
- A model can store a huge amount of knowledge in its total parameters.
- Each token only pays for the active parameters, so inference is faster and cheaper.
For example, Alibaba’s open-weights Qwen3.8-Max is reported at about 2.4 trillion total parameters with roughly 95 billion active per token, and DeepSeek’s V4 family also uses MoE. See our model tracker.
Trade-offs
MoE models need a lot of memory, because all experts must be loaded even if few are used. Training is harder to balance, because routers can overuse some experts. That’s why “you can download it” doesn’t always mean “you can run it” for the largest open-weights models.
History
The idea dates to 1991 research by Robert Jacobs, Michael Jordan, Steven Nowlan and Geoffrey Hinton. Google’s 2017 “sparsely-gated mixture-of-experts” paper brought it to large neural networks, and it became mainstream in frontier language models in the mid-2020s.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.