flâneur — a map of the web's best reading

Untitled

alignment.anthropic.com · 2,570 words · saved by 1 readers

We release AuditBench, an alignment auditing benchmark. AuditBench consists of 56 language models with implanted hidden behaviors—such as sycophantic deference, opposition to AI regulation, or hidden loyalties—which they do not confess to when asked. We also develop an agent that audits models using a configurable set of tools. Using this agent, we study which tools are most effective for auditing. Alignment auditing—investigating AI systems to uncover hidden or unintended behaviors—is a core challenge for safe AI deployment. Recent work has explored automated investigator agents: language model agents equipped with various tools that can probe a target model for problematic behavior. But, basic questions about investigator agents remain wide open: Which tools are actually worth using? Which agent scaffolds work best? How should tools be structured to maximize their value? Progress on these questions has been bottlenecked by a lack of standardized testbed for evaluating investigator ag

AuditBench Alignment Science Blog AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors Abhay Sheshadri March 10, 2026 Aidan Ewart, Kai Fronsdal, Isha Gupta Samuel R. Bowman, Sara Price, Samuel Marks, Rowan Wang tl;dr We release AuditBench, an alignment auditing benchmark. AuditBench consists of 56 language models with implanted hidden behaviors—such as sycophantic deference, opposition to AI regulation, or hidden loyalties—which they do not confess to when asked. We also develop an agent that audits models using a configurable set of tools. Using this agent, we

Explore this link on the map →

related reading