flâneur — a map of the web's best reading

Auditing language models for hidden objectives \ Anthropic

anthropic.com · 2,911 words · saved by 1 readers

A collaboration between Anthropic's Alignment Science and Interpretability teams

Alignment Interpretability Auditing language models for hidden objectives Mar 13, 2025 Read the paper A new paper from the Anthropic Alignment Science and Interpretability teams studies alignment audits —systematic investigations into whether models are pursuing hidden objectives. We practice alignment audits by deliberately training a language model with a hidden misaligned objective and asking teams of blinded researchers to investigate it. This exercise built practical experience conducting alignment audits and served as a testbed for developing auditing techniques for future study. In King

Explore this link on the map →

related reading