SI Glossary · Safety & alignment
Interpretability
On this page
Modern SI models are often called “black boxes”: we know their inputs, outputs and billions of parameters, but not why they do what they do. Interpretability research tries to open the box.
Two main approaches
- Explainability (post-hoc): methods that explain individual decisions, such as which pixels or words most influenced a prediction. These are common in regulated settings like lending.
- Mechanistic interpretability: reverse-engineering the internal circuits of a neural network, identifying human-understandable “features” (concepts like the Golden Gate Bridge or deception) and tracing how they combine into behaviour.
Notable progress
Labs including Anthropic, Google DeepMind and OpenAI have used techniques such as sparse autoencoders to find millions of interpretable features inside production models, and “circuit tracing” to follow a model’s internal steps on specific tasks. In demonstrations, researchers have amplified or suppressed features to change a model’s behaviour.
Why it matters
Interpretability could let developers detect deception, hidden goals or dangerous knowledge before deployment, a central hope for alignment of more capable systems. Regulators and auditors, including the independent auditors envisioned in the SI Accord, may increasingly rely on such tools.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.