SI Glossary · Core concepts
Foundation Model
On this page
A foundation model is a large model trained once, at great expense, on broad data, typically text, code and images from the internet and licensed sources, and then adapted for many uses. The term was coined by Stanford researchers in 2021 to describe models like GPT-3 and BERT that serve as the “foundation” for countless applications.
How foundation models are used
- Pretraining on a vast dataset teaches general knowledge and skills.
- Adaptation, through fine-tuning, RLHF, prompting or retrieval, shapes the model for a specific product.
- Deployment happens via apps, APIs or open weights.
In law and policy
Regulators use related terms. The EU AI Act regulates general-purpose AI (GPAI) models and applies extra duties to those with “systemic risk”. U.S. state laws such as California’s SB 53 use the term “foundation model” and add obligations for the largest “frontier” developers. See frontier model and compute threshold.
Examples
The GPT, Claude, Gemini, Llama, Grok, Qwen, DeepSeek and Mistral model families are all foundation models. See our frontier model tracker.
Sources
- On the Opportunities and Risks of Foundation Models — Bommasani et al., Stanford CRFM, 2021
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.