SI Glossary · Models & architecture
Context Window
On this page
A model’s context window is its working memory: everything it can “see” when producing a response. That includes the conversation so far, any files or web pages you’ve provided, hidden instructions from the app, and the answer it’s writing. It’s measured in tokens.
How big are context windows?
| Year | Typical frontier context |
|---|---|
| 2020 (GPT-3) | ~2,000 tokens |
| 2023 (GPT-4) | 8,000–32,000 tokens |
| 2024 | 128,000–1,000,000 tokens |
| 2026 | ~1,000,000 tokens is standard at the frontier |
As of October 2026, OpenAI’s GPT-6 models offer about 1.05 million tokens, Anthropic’s Claude Fable 5.1, Opus 5.5 and Sonnet 5.5 offer 1 million, and xAI’s Grok 4.7 offers 500,000. Compare them in the model tracker.
Bigger isn’t automatically better
- Cost and speed: you pay for every token you send, and long inputs are slower.
- “Lost in the middle”: models can overlook details buried in very long inputs, though this has improved.
- Alternatives: retrieval-augmented generation (RAG) fetches only the relevant passages instead of sending everything.
Context window vs output limit
The maximum output is separate and smaller. Many 2026 frontier models cap a single reply at around 128,000 tokens, while Google says Gemini 4 Argon can produce up to a million.
Frequently asked questions
How big is a 1-million-token context window?
Roughly 500,000 to 750,000 English words, depending on the tokenizer: several long novels or a large codebase at once.
What happens when you exceed the context window?
The model can't see the overflow. Apps either refuse, cut off older parts of the conversation, or summarize them. Retrieval-augmented generation is another way to work with more information than fits.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.