Skip to content
SI.info

SI Glossary · Safety & alignment

Interpretability

Published 1 min read
On this page
  1. Two main approaches
  2. Notable progress
  3. Why it matters

Modern SI models are often called “black boxes”: we know their inputs, outputs and billions of parameters, but not why they do what they do. Interpretability research tries to open the box.

Two main approaches

  • Explainability (post-hoc): methods that explain individual decisions, such as which pixels or words most influenced a prediction. These are common in regulated settings like lending.
  • Mechanistic interpretability: reverse-engineering the internal circuits of a neural network, identifying human-understandable “features” (concepts like the Golden Gate Bridge or deception) and tracing how they combine into behaviour.

Notable progress

Labs including Anthropic, Google DeepMind and OpenAI have used techniques such as sparse autoencoders to find millions of interpretable features inside production models, and “circuit tracing” to follow a model’s internal steps on specific tasks. In demonstrations, researchers have amplified or suppressed features to change a model’s behaviour.

Why it matters

Interpretability could let developers detect deception, hidden goals or dangerous knowledge before deployment, a central hope for alignment of more capable systems. Regulators and auditors, including the independent auditors envisioned in the SI Accord, may increasingly rely on such tools.

← Back to the SI Glossary

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy