SI Glossary · Models & architecture
Parameters
A model’s parameters, also called its weights, are the billions of numbers that encode everything it has learned. During training, each parameter is nudged up or down to make the model’s predictions more accurate. After training, the parameters are the model: share them and you’ve shared the model. That’s what open-weights releases do.
How big are models?
| Model | Year | Parameters |
|---|---|---|
| GPT-2 | 2019 | 1.5 billion |
| GPT-3 | 2020 | 175 billion |
| DeepSeek V4 Pro | 2026 | ~1.6–1.7 trillion total (reported), MoE |
| Qwen3.8-Max | 2026 | ~2.4 trillion total, ~95 billion active (reported) |
Many frontier labs, including OpenAI, Anthropic and Google, no longer disclose parameter counts.
Total vs active parameters
In mixture-of-experts models, only some parameters are used for each token. “Active parameters” better predicts running cost, while “total parameters” relates to stored knowledge and memory needs.
More parameters, smarter model?
Usually, up to a point. Scaling laws show performance improves predictably with more parameters, data and compute together. But data quality, training methods and post-training matter just as much, and a well-trained smaller model can beat a poorly trained larger one.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.