Auditing language models for hidden objectives \ Anthropic
anthropic.com · 2,911 words · saved by 1 readers
A collaboration between Anthropic's Alignment Science and Interpretability teams
Alignment Interpretability Auditing language models for hidden objectives Mar 13, 2025 Read the paper A new paper from the Anthropic Alignment Science and Interpretability teams studies alignment audits —systematic investigations into whether models are pursuing hidden objectives. We practice alignment audits by deliberately training a language model with a hidden misaligned objective and asking teams of blinded researchers to investigate it. This exercise built practical experience conducting alignment audits and served as a testbed for developing auditing techniques for future study. In King
related reading
- Auditing language models for hidden objectives — LessWronglesswrong.com
- [2503.10965] Auditing language models for hidden objectivesarxiv.org
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- AuditBenchalignment.anthropic.com
- Building and evaluating alignment auditing agents — AI Alignment Forumalignmentforum.org
- Building and evaluating alignment auditing agentsalignment.anthropic.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Pre-deployment auditing can catch an overt saboteuralignment.anthropic.com
- Alignment Faking Mitigationsalignment.anthropic.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org