flâneur — a map of the web's best reading

Auditing language models for hidden objectives — LessWrong

lesswrong.com · 5,466 words · saved by 1 readers

We study alignment audits—systematic investigations into whether an AI is pursuing hidden objectives—by training a model with a hidden misaligned obj…

x Auditing language models for hidden objectives — LessWrong AI Auditing Interpretability (ML & AI) AI Frontpage 2025 Top Fifty: 9 % 153 Auditing language models for hidden objectives by Sam Marks , Johannes Treutlein , dmz , Sam Bowman , Hoagy , Carson Denison , Kei Nishimura-Gasparian , 7vik , Akbir Khan , Austin Meek , Euan Ong , Christopher Olah , Fabien Roger , jeanne_ , Meg , Drake Thomas , Adam Jermyn , Monte M , evhub 13th Mar 2025 AI Alignment Forum 15 min read 15 153 Ω 84 We study alignment audits —systematic investigations into whether an AI is pursuing hidden objectives—by training

Explore this link on the map →

saved by

related reading