SI Glossary · Safety & alignment
Red Teaming
On this page
Red teaming borrows a term from military and cybersecurity practice: a “red team” plays the adversary. For SI, red teamers try to break a model: get it to give instructions for weapons, leak private data, write malware, deceive users or ignore its rules.
Who does it
- Internal teams at each lab
- External experts in biology, chemistry, cybersecurity and other fields, hired before launch
- Government testing bodies, such as the UK AI Security Institute and the U.S. Center for AI Standards and Innovation (CAISI)
- The public, through bug-bounty programmes and events
Common techniques
- Jailbreaks: prompts designed to bypass safety training
- Prompt injection: hiding malicious instructions in content an agent reads
- Capability elicitation: testing whether a model can meaningfully assist in serious harm
- Automated red teaming: using other models to generate attacks at scale
Where results go
Findings feed into safety fixes and are summarised in model cards or system cards. Under laws like California SB 53, large developers must publish safety frameworks describing how they assess catastrophic risks.
Written by
Luka Kušec · Editor
Editor of SI.info. Writes about Super Intelligence, technology policy and the people building frontier models.