Pre-deployment auditing can catch an overt saboteur
We test whether our production pre-deployment alignment auditing methods can catch overt saboteurs: models trained to sabotage Anthropic but not to conceal this behavior. A human working together with an automated auditing agent successfully identified three overt sabotage agents, without flagging two benign models. This suggests that our pre-deployment alignment audit would detect if current Claude models were overtly trying to sabotage Anthropic. Anthropic uses existing Claude models to assist in the development of future generations of Claude models. This includes generating training data, writing research code, monitoring actions taken by AI agents on Anthropic infrastructure, and assessing the alignment of trained models. If existing models behave reliably, then using them to assist with safety R&D allows us to carry out this work more effectively. However, using AI for AI development also presents risks. In the worst case, Claude models might actively undermine Anthropic’s alignm
Johannes Treutlein, Samuel R. Bowman, Trenton Bricken January 28, 2026 Alex Cloud, Misha Wagner, Rowan Wang Evan Hubinger, Fabien Roger, Sam Marks tl;dr We test whether our production pre-deployment alignment auditing methods can catch overt saboteurs: models trained to sabotage Anthropic but not to conceal this behavior. A human working together with an automated auditing agent successfully identified three overt sabotage agents, without flagging two benign models. This suggests that our pre-deployment alignment audit would detect if current Claude models were overtly trying to…
saved by
related reading
- Thoughts on Claude Fable's silent safeguards — LessWronglesswrong.com
- Teaching Claude Whyalignment.anthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Claude 4 System Cardwww-cdn.anthropic.com
- Auditing language models for hidden objectives — LessWronglesswrong.com
- Building and evaluating alignment auditing agents — AI Alignment Forumalignmentforum.org
- An alignment assessment of recent cybersecurity incidentsanthropic.com