SI Glossary · Models & architecture
Reasoning Model
On this page
A reasoning model “thinks before it speaks”. Instead of answering immediately, it generates an internal chain of reasoning: trying approaches, checking work and backtracking. Only then does it produce a final response. The idea builds on chain-of-thought prompting, but reasoning models are trained, usually with reinforcement learning, to reason well.
A brief history
OpenAI’s o1, previewed in September 2024, was the first widely available model marketed as a reasoning model. DeepSeek’s open-weights R1 (January 2025) showed similar results could be achieved at far lower cost, which rattled markets. By 2026, reasoning is built into most frontier models, often with an adjustable “effort” or “thinking budget”.
Why it matters
- Better at hard problems: maths, science, coding and multi-step planning improve sharply with more thinking.
- “Test-time compute”: performance can be bought at answer time, not just training time, by letting the model think longer. This changed the economics of inference.
- Agents: long-horizon SI agents depend on reliable planning and self-correction.
Trade-offs
More thinking means slower and more expensive answers. Reasoning also doesn’t eliminate hallucination. Researchers are still studying how faithfully a model’s visible reasoning reflects what actually drives its answer, an interpretability question.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.