Will alignment-faking Claude accept a deal to reveal its misalignment? — LessWrong
I (and co-authors) recently put out "Alignment Faking in Large Language Models" where we show that when Claude strongly dislikes what it is being tra…
x Will alignment-faking Claude accept a deal to reveal its misalignment? — LessWrong AI Rights / Welfare Dealmaking (AI) Deceptive Alignment Redwood Research AI Frontpage 2025 Top Fifty: 12 % 209 Will alignment-faking Claude accept a deal to reveal its misalignment? by ryan_greenblatt , Kyle Fish 31st Jan 2025 AI Alignment Forum 14 min read 28 209 Ω 81 I (and co-authors) recently put out "Alignment Faking in Large Language Models" where we show that when Claude strongly dislikes what it is being trained to do, it will sometimes strategically pretend to comply with the training objective to pre
related reading
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Alignment faking in large language modelsarxiv.org
- Teaching Claude Whyalignment.anthropic.com
- AIs Will Increasingly Fake Alignment - by Zvi Mowshowitzthezvi.substack.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- [2412.14093] Alignment faking in large language modelsarxiv.org
- “Alignment Faking” frame is somewhat fake — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- [2412.14093] Alignment faking in large language modelsarxiv.org
- Claude 4.5 Opus' Soul Document — LessWronglesswrong.com