Skip to content
SI.info

SI Glossary · Using SI

Benchmark

Published 1 min read
On this page
  1. Well-known benchmarks
  2. Why benchmarks are tricky
  3. Our advice
  4. Safety evaluations

A benchmark is an exam for SI models: a fixed set of tasks with known answers, so different models can be scored the same way. Labs publish benchmark results with every major release.

Well-known benchmarks

BenchmarkTests
MMLUKnowledge across 57 academic subjects
GPQAVery hard, “Google-proof” science questions
SWE-benchFixing real bugs in open-source software
Humanity’s Last ExamExpert-level questions across fields
ARC-AGINovel visual puzzles that test general reasoning
Chatbot ArenaHuman votes in blind head-to-head comparisons

Why benchmarks are tricky

  • Saturation: frontier models quickly max out older tests, so new, harder ones keep appearing.
  • Contamination: test questions can leak into training data, inflating scores.
  • Different settings: labs report results with different amounts of thinking time, tools or attempts.
  • Real world ≠ test: a high score doesn’t guarantee reliability on your tasks.

Our advice

Use benchmarks to shortlist models, then test them on 20–50 of your own real tasks. Our model comparisons focus on verifiable facts (price, context window, availability) rather than marketing benchmark claims.

Safety evaluations

Labs also run dangerous capability evaluations for cyber-offence, biology and autonomy. These underpin safety frameworks such as responsible scaling policies and are complemented by red teaming.

← Back to the SI Glossary

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy