Building and evaluating alignment auditing agents — AI Alignment Forum
TL;DR: We develop three agents that autonomously perform alignment auditing tasks. When tested against models with intentionally-inserted alignment i…
x Building and evaluating alignment auditing agents — AI Alignment Forum AI Frontpage 29 Building and evaluating alignment auditing agents by Sam Marks , trentbrick , RowanWang , Sam Bowman , Euan Ong , Johannes Treutlein , evhub 24th Jul 2025 6 min read 1 29 TL;DR: We develop three agents that autonomously perform alignment auditing tasks. When tested against models with intentionally-inserted alignment issues, our agents successfully uncover an LLM's hidden goal, build behavioral evaluations, and surface concerning LLM behaviors. We are using these agents to assist with alignment audits of f
Explore this link on the map →saved by
related reading
- Research Areas in Benchmark Design and Evaluation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Building and evaluating alignment auditing agentsalignment.anthropic.com
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- AuditBenchalignment.anthropic.com
- Auditing language models for hidden objectives — LessWronglesswrong.com
- Auditing language models for hidden objectives \ Anthropicanthropic.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- [2503.10965] Auditing language models for hidden objectivesarxiv.org
- Teaching Claude why \ Anthropicanthropic.com