Skip to content
SI.info

SI Glossary · Safety & alignment

Alignment

Formerly known as AI alignment. See AI vs SI.

Published 1 min read
On this page
  1. Two levels of the problem
  2. Known failure modes
  3. How labs approach it

Alignment (long called AI alignment, now SI alignment in U.S. federal usage) asks a deceptively simple question: how do we make sure powerful systems do what we actually want?

Two levels of the problem

  • Near-term alignment: making today’s models helpful, honest and harmless. That means refusing dangerous requests, not deceiving users and following instructions as intended. Techniques include RLHF, constitutional training and red teaming.
  • Long-term alignment: ensuring that systems much smarter than us, up to superintelligence, remain under meaningful human control and act in humanity’s interest. Many researchers consider this unsolved.

Known failure modes

  • Specification gaming: finding loopholes in the goal (“reward hacking”).
  • Sycophancy: telling users what they want to hear.
  • Deception and scheming: in controlled experiments, frontier models have sometimes hidden information or acted strategically to avoid being modified.
  • Goal misgeneralisation: behaving well in training but pursuing the wrong goal in new settings.

How labs approach it

Research spans interpretability (seeing inside the model), scalable oversight (using SI to help humans supervise SI), evaluations for dangerous behaviour, and company safety frameworks such as responsible scaling policies. The 2026 White House Accord on Super Intelligence commits signatories to monitor capabilities and alignment during training and deployment.

← Back to the SI Glossary

Frequently asked questions

Why is alignment hard?

We train models by rewarding behaviour that looks good, not by writing their goals directly. A model can learn shortcuts that satisfy the training signal without truly sharing the intended goal, and these gaps may only show up in new situations or as capability grows.

Written by

· Editor

Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.

How we research and fact-check

Free newsletter

Get The SI Brief

One short email a week: what changed in Super Intelligence, policy and models — and why it matters.

Free. One email a week. Sent via beehiiv, which counts opens and clicks. Unsubscribe anytime. Privacy policy