flâneur — a map of the web's best reading

Will alignment-faking Claude accept a deal to reveal its misalignment? — LessWrong

lesswrong.com · 12,449 words · saved by 1 readers

I (and co-authors) recently put out "Alignment Faking in Large Language Models" where we show that when Claude strongly dislikes what it is being tra…

x Will alignment-faking Claude accept a deal to reveal its misalignment? — LessWrong AI Rights / Welfare Dealmaking (AI) Deceptive Alignment Redwood Research AI Frontpage 2025 Top Fifty: 12 % 209 Will alignment-faking Claude accept a deal to reveal its misalignment? by ryan_greenblatt , Kyle Fish 31st Jan 2025 AI Alignment Forum 14 min read 28 209 Ω 81 I (and co-authors) recently put out "Alignment Faking in Large Language Models" where we show that when Claude strongly dislikes what it is being trained to do, it will sometimes strategically pretend to comply with the training objective to pre

Explore this link on the map →

related reading