Untitled
We release AuditBench, an alignment auditing benchmark. AuditBench consists of 56 language models with implanted hidden behaviors—such as sycophantic deference, opposition to AI regulation, or hidden loyalties—which they do not confess to when asked. We also develop an agent that audits models using a configurable set of tools. Using this agent, we study which tools are most effective for auditing. Alignment auditing—investigating AI systems to uncover hidden or unintended behaviors—is a core challenge for safe AI deployment. Recent work has explored automated investigator agents: language model agents equipped with various tools that can probe a target model for problematic behavior. But, basic questions about investigator agents remain wide open: Which tools are actually worth using? Which agent scaffolds work best? How should tools be structured to maximize their value? Progress on these questions has been bottlenecked by a lack of standardized testbed for evaluating investigator ag
AuditBench Alignment Science Blog AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors Abhay Sheshadri March 10, 2026 Aidan Ewart, Kai Fronsdal, Isha Gupta Samuel R. Bowman, Sara Price, Samuel Marks, Rowan Wang tl;dr We release AuditBench, an alignment auditing benchmark. AuditBench consists of 56 language models with implanted hidden behaviors—such as sycophantic deference, opposition to AI regulation, or hidden loyalties—which they do not confess to when asked. We also develop an agent that audits models using a configurable set of tools. Using this agent, we
Explore this link on the map →related reading
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- Building and evaluating alignment auditing agents — AI Alignment Forumalignmentforum.org
- Auditing language models for hidden objectives — LessWronglesswrong.com
- Building and evaluating alignment auditing agentsalignment.anthropic.com
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Auditing language models for hidden objectives \ Anthropicanthropic.com
- [2503.10965] Auditing language models for hidden objectivesarxiv.org
- Natural Language Autoencoders \ Anthropicanthropic.com
- How confessions can keep language models honest | OpenAIopenai.com
- Petri: An open-source auditing tool to accelerate AI safety researchalignment.anthropic.com