✳flâneur — a map of the web's best reading
Auditing language models for hidden objectives \ Anthropic
anthropic.com · 2,911 words · saved by 1 readers
A collaboration between Anthropic's Alignment Science and Interpretability teams
Alignment Interpretability Auditing language models for hidden objectives Mar 13, 2025 Read the paper A new paper from the Anthropic Alignment Science and Interpretability teams studies alignment audits —systematic investigations into whether models are pursuing hidden objectives. We practice alignment audits by deliberately training a language model with a hidden misaligned objective and asking teams of blinded researchers to investigate it. This exercise built practical experience conducting alignment audits and served as a testbed for developing auditing techniques for future study. In King
Explore this link on the map →related reading
- Auditing language models for hidden objectives — LessWronglesswrong.com
- [2503.10965] Auditing language models for hidden objectivesarxiv.org
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- Building and evaluating alignment auditing agents — AI Alignment Forumalignmentforum.org
- Natural Language Autoencoders \ Anthropicanthropic.com
- AuditBenchalignment.anthropic.com
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- Building and evaluating alignment auditing agentsalignment.anthropic.com
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How confessions can keep language models honest | OpenAIopenai.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org