Skip to content
SI.info

SI Glossary · Models & architecture

Parameters

Published 1 min read
On this page
  1. How big are models?
  2. Total vs active parameters
  3. More parameters, smarter model?

A model’s parameters, also called its weights, are the billions of numbers that encode everything it has learned. During training, each parameter is nudged up or down to make the model’s predictions more accurate. After training, the parameters are the model: share them and you’ve shared the model. That’s what open-weights releases do.

How big are models?

ModelYearParameters
GPT-220191.5 billion
GPT-32020175 billion
DeepSeek V4 Pro2026~1.6–1.7 trillion total (reported), MoE
Qwen3.8-Max2026~2.4 trillion total, ~95 billion active (reported)

Many frontier labs, including OpenAI, Anthropic and Google, no longer disclose parameter counts.

Total vs active parameters

In mixture-of-experts models, only some parameters are used for each token. “Active parameters” better predicts running cost, while “total parameters” relates to stored knowledge and memory needs.

More parameters, smarter model?

Usually, up to a point. Scaling laws show performance improves predictably with more parameters, data and compute together. But data quality, training methods and post-training matter just as much, and a well-trained smaller model can beat a poorly trained larger one.

← Back to the SI Glossary

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy