Will alignment-faking Claude accept a deal to reveal its misalignment? — LessWrong
I (and co-authors) recently put out "Alignment Faking in Large Language Models" where we show that when Claude strongly dislikes what it is being tra…
x Will alignment-faking Claude accept a deal to reveal its misalignment? — LessWrong AI Rights / Welfare Dealmaking (AI) Deceptive Alignment Redwood Research AI Frontpage 2025 Top Fifty: 12 % 209 Will alignment-faking Claude accept a deal to reveal its misalignment? by ryan_greenblatt , Kyle Fish 31st Jan 2025 AI Alignment Forum 14 min read 28 209 Ω 81 I (and co-authors) recently put out "Alignment Faking in Large Language Models" where we show that when Claude strongly dislikes what it is being trained to do, it will sometimes strategically pretend to comply with the training objective to pre
Explore this link on the map →related reading
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Alignment faking in large language modelsarxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AIs Will Increasingly Fake Alignment - by Zvi Mowshowitzthezvi.substack.com
- Teaching Claude why \ Anthropicanthropic.com
- “Alignment Faking” frame is somewhat fake — LessWronglesswrong.com
- [2412.14093] Alignment faking in large language modelsarxiv.org
- [2412.14093] Alignment faking in large language modelsarxiv.org
- Alignment Faking Mitigationsalignment.anthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Claude’s Character \ Anthropicanthropic.com