Untitled
We release AuditBench, an alignment auditing benchmark. AuditBench consists of 56 language models with implanted hidden behaviors—such as sycophantic deference, opposition to AI regulation, or hidden loyalties—which they do not confess to when asked. We also develop an agent that audits models using a configurable set of tools. Using this agent, we study which tools are most effective for auditing. Alignment auditing—investigating AI systems to uncover hidden or unintended behaviors—is a core challenge for safe AI deployment. Recent work has explored automated investigator agents: language model agents equipped with various tools that can probe a target model for problematic behavior. But, basic questions about investigator agents remain wide open: Which tools are actually worth using? Which agent scaffolds work best? How should tools be structured to maximize their value? Progress on these questions has been bottlenecked by a lack of standardized testbed for evaluating investigator ag
AuditBench Alignment Science Blog AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors Abhay Sheshadri March 10, 2026 Aidan Ewart, Kai Fronsdal, Isha Gupta Samuel R. Bowman, Sara Price, Samuel Marks, Rowan Wang tl;dr We release AuditBench, an alignment auditing benchmark. AuditBench consists of 56 language models with implanted hidden behaviors—such as sycophantic deference, opposition to AI regulation, or hidden loyalties—which they do not confess to when asked. We also develop an agent that audits models using a configurable set of tools. Using this agent, we
related reading
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- Auditing language models for hidden objectives — LessWronglesswrong.com
- Building and evaluating alignment auditing agents — AI Alignment Forumalignmentforum.org
- Building and evaluating alignment auditing agentsalignment.anthropic.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Auditing language models for hidden objectives \ Anthropicanthropic.com
- Pre-deployment auditing can catch an overt saboteuralignment.anthropic.com
- [2503.10965] Auditing language models for hidden objectivesarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org