Auditing language models for hidden objectives — LessWrong
We study alignment audits—systematic investigations into whether an AI is pursuing hidden objectives—by training a model with a hidden misaligned obj…
x Auditing language models for hidden objectives — LessWrong AI Auditing Interpretability (ML & AI) AI Frontpage 2025 Top Fifty: 9 % 153 Auditing language models for hidden objectives by Sam Marks , Johannes Treutlein , dmz , Sam Bowman , Hoagy , Carson Denison , Kei Nishimura-Gasparian , 7vik , Akbir Khan , Austin Meek , Euan Ong , Christopher Olah , Fabien Roger , jeanne_ , Meg , Drake Thomas , Adam Jermyn , Monte M , evhub 13th Mar 2025 AI Alignment Forum 15 min read 15 153 Ω 84 We study alignment audits —systematic investigations into whether an AI is pursuing hidden objectives—by training
saved by
related reading
- [2503.10965] Auditing language models for hidden objectivesarxiv.org
- Auditing language models for hidden objectives \ Anthropicanthropic.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- AuditBenchalignment.anthropic.com
- Building and evaluating alignment auditing agents — AI Alignment Forumalignmentforum.org
- Building and evaluating alignment auditing agentsalignment.anthropic.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Pre-deployment auditing can catch an overt saboteuralignment.anthropic.com
- Alignment Faking Mitigationsalignment.anthropic.com