SI Glossary · Safety & alignment
Responsible Scaling Policy (RSP)
On this page
A responsible scaling policy (RSP) is an “if–then” safety commitment. If a model shows certain dangerous capabilities in testing, then the developer applies stronger security and deployment safeguards before going further, or stops if it can’t.
Who has one
- Anthropic: Responsible Scaling Policy, first published September 2023, with “AI Safety Levels” (ASL-1 to ASL-4+) modelled on biosafety levels.
- OpenAI: Preparedness Framework, with risk categories such as cybersecurity and biology and thresholds that trigger safeguards. GPT-6 Astra was reported as the first model to reach its “Critical” cybersecurity threshold.
- Google DeepMind: Frontier Safety Framework, with “critical capability levels”.
- Many other developers signed on to publishing such frameworks at the 2024 AI Seoul Summit.
From voluntary to required
What began as voluntary policy is becoming law in some places. California SB 53 (in effect January 2026) requires large frontier developers to publish a safety framework and report critical incidents. New York’s RAISE Act does similar from 2027. The 2026 White House Accord on Super Intelligence adds independent auditors and board oversight on a voluntary basis.
Criticisms
Critics note that companies write and can revise their own thresholds, and that commercial pressure can push toward lenient interpretations. Supporters say RSPs create transparency and accountability that didn’t exist before.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.