[2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors
Abstract:We introduce AuditBench, an alignment auditing benchmark. AuditBench consists of 56 language models with implanted hidden behaviors. Each model has one of 14 concerning behaviors--such as sycophantic deference, opposition to AI regulation, or secret geopolitical loyalties--which it does not confess to when directly asked. AuditBench models are highly diverse--some are subtle, while others are overt, and we use varying training techniques both for implanting behaviors and training models not to confess. To demonstrate AuditBench's utility, we develop an investigator agent that autonomously employs a configurable set of auditing tools. By measuring investigator agent success using different tools, we can evaluate their efficacy. Notably, we observe a tool-to-agent gap, where tools that perform well in standalone non-agentic evaluations fail to translate into improved performance when used with our investigator agent. We find that our most effective tools involve scaffolded calls to auxiliary models that generate diverse prompts for the target. White-box interpretability tools can be helpful, but the agent performs best with black-box tools. We also find that audit success varies greatly across training techniques: models trained on synthetic documents are easier to audit than models trained on demonstrations, with better adversarial training further increasing auditing difficulty. We release our models, agent, and evaluation framework to support future quantitative, iterative science on alignment auditing.
[2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors --> Computer Science > Computation and Language arXiv:2602.22755 (cs) [Submitted on 26 Feb 2026 ( v1 ), last revised 9 Mar 2026 (this version, v3)] Title: AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors Authors: Abhay Sheshadri , Aidan Ewart , Kai Fronsdal , Isha Gupta , Samuel R. Bowman , Sara Price , Samuel Marks , Rowan Wang View a PDF of the paper titled AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors, by Abhay Sheshadri
Explore this link on the map →saved by
related reading
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- AuditBenchalignment.anthropic.com
- Auditing language models for hidden objectives — LessWronglesswrong.com
- Building and evaluating alignment auditing agents — AI Alignment Forumalignmentforum.org
- [2503.10965] Auditing language models for hidden objectivesarxiv.org
- Building and evaluating alignment auditing agentsalignment.anthropic.com
- Auditing language models for hidden objectives \ Anthropicanthropic.com
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- secret-loyalties-whitepaper.pdfformationresearch.com