flâneur

Pre-deployment auditing can catch an overt saboteur

alignment.anthropic.com · 5,924 words · saved by 2 readers

We test whether our production pre-deployment alignment auditing methods can catch overt saboteurs: models trained to sabotage Anthropic but not to conceal this behavior. A human working together with an automated auditing agent successfully identified three overt sabotage agents, without flagging two benign models. This suggests that our pre-deployment alignment audit would detect if current Claude models were overtly trying to sabotage Anthropic. Anthropic uses existing Claude models to assist in the development of future generations of Claude models. This includes generating training data, writing research code, monitoring actions taken by AI agents on Anthropic infrastructure, and assessing the alignment of trained models. If existing models behave reliably, then using them to assist with safety R&D allows us to carry out this work more effectively. However, using AI for AI development also presents risks. In the worst case, Claude models might actively undermine Anthropic’s alignm

Johannes Treutlein, Samuel R. Bowman, Trenton Bricken January 28, 2026 Alex Cloud, Misha Wagner, Rowan Wang Evan Hubinger, Fabien Roger, Sam Marks tl;dr We test whether our production pre-deployment alignment auditing methods can catch overt saboteurs: models trained to sabotage Anthropic but not to conceal this behavior. A human working together with an automated auditing agent successfully identified three overt sabotage agents, without flagging two benign models. This suggests that our pre-deployment alignment audit would detect if current Claude models were overtly trying to…

saved by

related reading