SI Glossary · Using SI
Benchmark
A benchmark is an exam for SI models: a fixed set of tasks with known answers, so different models can be scored the same way. Labs publish benchmark results with every major release.
Well-known benchmarks
| Benchmark | Tests |
|---|---|
| MMLU | Knowledge across 57 academic subjects |
| GPQA | Very hard, “Google-proof” science questions |
| SWE-bench | Fixing real bugs in open-source software |
| Humanity’s Last Exam | Expert-level questions across fields |
| ARC-AGI | Novel visual puzzles that test general reasoning |
| Chatbot Arena | Human votes in blind head-to-head comparisons |
Why benchmarks are tricky
- Saturation: frontier models quickly max out older tests, so new, harder ones keep appearing.
- Contamination: test questions can leak into training data, inflating scores.
- Different settings: labs report results with different amounts of thinking time, tools or attempts.
- Real world ≠ test: a high score doesn’t guarantee reliability on your tasks.
Our advice
Use benchmarks to shortlist models, then test them on 20–50 of your own real tasks. Our model comparisons focus on verifiable facts (price, context window, availability) rather than marketing benchmark claims.
Safety evaluations
Labs also run dangerous capability evaluations for cyber-offence, biology and autonomy. These underpin safety frameworks such as responsible scaling policies and are complemented by red teaming.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.