AIs Will Increasingly Fake Alignment - by Zvi Mowshowitz
This post goes over the important and excellent new paper from Anthropic and Redwood Research, with Ryan Greenblatt as lead author, Alignment Faking in Large Language Models.
AIs Will Increasingly Fake Alignment Zvi Mowshowitz Dec 24, 2024 43 17 6 Share This post goes over the important and excellent new paper from Anthropic and Redwood Research, with Ryan Greenblatt as lead author, Alignment Faking in Large Language Models. This is by far the best demonstration so far of the principle that AIs Will Increasingly Attempt Shenanigans . This was their announcement thread. New Anthropic research: Alignment faking in large language models. In a series of experiments with Redwood Research, we found that Claude often pretends to have different views during training, while
Explore this link on the map →related reading
- Alignment faking in large language modelsarxiv.org
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- “Alignment Faking” frame is somewhat fake — LessWronglesswrong.com
- [2412.14093] Alignment faking in large language modelsarxiv.org
- [2412.14093] Alignment faking in large language modelsarxiv.org
- Alignment Faking Mitigationsalignment.anthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Will alignment-faking Claude accept a deal to reveal its misalignment? — LessWronglesswrong.com
- Towards training-time mitigations for alignment faking in RL — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Alignment will happen by default. What’s next? — LessWronglesswrong.com