flâneur — a map of the web's best reading

Building and evaluating alignment auditing agents

alignment.anthropic.com · 13,381 words · saved by 1 readers

We develop three agents that autonomously perform alignment auditing tasks. When tested against models with intentionally-inserted alignment issues, our agents successfully uncover an LLM's hidden goal, build behavioral evaluations, and surface concerning LLM behaviors. We are using these agents to assist with alignment audits of frontier models like Claude 4. As part of our pre-deployment testing of Claude Opus 4, we conducted an alignment audit: a systematic investigation into the model’s alignment-relevant propensities including sycophancy, alignment faking, hidden goals. Through this audit, we believe that Anthropic, and the public, gained a deeper understanding of Claude Opus 4's alignment than we've achieved for any previous model. However, alignment audits face two major challenges. The first major challenge is scalability: Alignment audits require a large investment of human researcher time. As the pace of AI development accelerates—resulting in more models to audit and more gr

Building and evaluating alignment auditing agents Alignment Science Blog Building and evaluating alignment auditing agents Trenton Bricken, Rowan Wang, Sam Bowman July 24, 2025 Euan Ong, Johannes Treutlein, Jeff Wu Evan Hubinger, Samuel Marks We develop three agents that autonomously perform alignment auditing tasks. When tested against models with intentionally-inserted alignment issues, our agents successfully uncover an LLM's hidden goal, build behavioral evaluations, and surface concerning LLM behaviors. We are using these agents to assist with alignment audits of frontier models like Clau

Explore this link on the map →

related reading