“Alignment Faking” frame is somewhat fake — LessWrong
I like the research. I mostly trust the results. I dislike the 'Alignment Faking' name and frame, and I'm afraid it will stick and lead to more confu…
x “Alignment Faking” frame is somewhat fake — LessWrong Best of LessWrong 2024 Deceptive Alignment AI Frontpage 170 “Alignment Faking” frame is somewhat fake by Jan_Kulveit 20th Dec 2024 AI Alignment Forum 7 min read 16 170 Ω 73 I like the research. I mostly trust the results. I dislike the 'Alignment Faking' name and frame, and I'm afraid it will stick and lead to more confusion. This post offers a different frame. The main way I think about the result is: it's about capability - the model exhibits strategic preference preservation behavior ; also, harmlessness generalized better than honesty
Explore this link on the map →related reading
- AIs Will Increasingly Fake Alignment - by Zvi Mowshowitzthezvi.substack.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Alignment faking in large language modelsarxiv.org
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Will alignment-faking Claude accept a deal to reveal its misalignment? — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Alignment will happen by default. What’s next? — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Towards training-time mitigations for alignment faking in RL — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org