SI Glossary · Models & architecture
Large Language Model (LLM)
On this page
A large language model (LLM) is the technology behind chat assistants like ChatGPT, Claude, Gemini and Grok. At its core, an LLM does one thing: given some text, it predicts what comes next, one token at a time. Done at enormous scale, that simple task produces systems that can draft essays, write software, reason through problems and hold conversations.
How LLMs are built
- Pretraining: the model reads trillions of words of text and code and learns to predict the next token. This is where most of its knowledge comes from.
- Post-training: fine-tuning and RLHF teach it to follow instructions, be helpful and refuse harmful requests.
- Deployment: the model runs on GPUs or other chips, a process called inference, behind an app or API.
Almost all LLMs use the transformer architecture, introduced by Google researchers in 2017.
What “large” means
Size is measured in parameters, the adjustable numbers the model learns. GPT-2 (2019) had 1.5 billion. Frontier models in 2026 have hundreds of billions to several trillion, often using a mixture-of-experts design so only part of the model runs for each token.
Limits
- Hallucination: stating false information fluently.
- Knowledge cutoff: facts after the training date are unknown unless the model can search.
- Context window: a limit on how much text it can consider at once, though frontier models now handle around a million tokens.
Modern LLMs are often multimodal, reading images and audio as well as text, and many are reasoning models that “think” before answering.
Frequently asked questions
How does an LLM know things?
During pretraining it learns statistical patterns from trillions of words, which encode a great deal of factual and procedural knowledge. It doesn't look facts up unless it's connected to search or a database, which is why it can be confidently wrong.
Is ChatGPT an LLM?
ChatGPT is a product built on OpenAI's LLMs (the GPT family), with extra layers such as safety training, tools and memory. The same is true of Claude, Gemini and Grok.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.