Building and evaluating alignment auditing agents
We develop three agents that autonomously perform alignment auditing tasks. When tested against models with intentionally-inserted alignment issues, our agents successfully uncover an LLM's hidden goal, build behavioral evaluations, and surface concerning LLM behaviors. We are using these agents to assist with alignment audits of frontier models like Claude 4. As part of our pre-deployment testing of Claude Opus 4, we conducted an alignment audit: a systematic investigation into the model’s alignment-relevant propensities including sycophancy, alignment faking, hidden goals. Through this audit, we believe that Anthropic, and the public, gained a deeper understanding of Claude Opus 4's alignment than we've achieved for any previous model. However, alignment audits face two major challenges. The first major challenge is scalability: Alignment audits require a large investment of human researcher time. As the pace of AI development accelerates—resulting in more models to audit and more gr
Building and evaluating alignment auditing agents Alignment Science Blog Building and evaluating alignment auditing agents Trenton Bricken, Rowan Wang, Sam Bowman July 24, 2025 Euan Ong, Johannes Treutlein, Jeff Wu Evan Hubinger, Samuel Marks We develop three agents that autonomously perform alignment auditing tasks. When tested against models with intentionally-inserted alignment issues, our agents successfully uncover an LLM's hidden goal, build behavioral evaluations, and surface concerning LLM behaviors. We are using these agents to assist with alignment audits of frontier models like Clau
Explore this link on the map →related reading
- Building and evaluating alignment auditing agents — AI Alignment Forumalignmentforum.org
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- AuditBenchalignment.anthropic.com
- Auditing language models for hidden objectives — LessWronglesswrong.com
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- Auditing language models for hidden objectives \ Anthropicanthropic.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- [2503.10965] Auditing language models for hidden objectivesarxiv.org
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Teaching Claude why \ Anthropicanthropic.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org