AIs Will Increasingly Fake Alignment - by Zvi Mowshowitz
This post goes over the important and excellent new paper from Anthropic and Redwood Research, with Ryan Greenblatt as lead author, Alignment Faking in Large Language Models.
AIs Will Increasingly Fake Alignment Zvi Mowshowitz Dec 24, 2024 43 17 6 Share This post goes over the important and excellent new paper from Anthropic and Redwood Research, with Ryan Greenblatt as lead author, Alignment Faking in Large Language Models. This is by far the best demonstration so far of the principle that AIs Will Increasingly Attempt Shenanigans . This was their announcement thread. New Anthropic research: Alignment faking in large language models. In a series of experiments with Redwood Research, we found that Claude often pretends to have different views during training, while
related reading
- Alignment faking in large language modelsarxiv.org
- Alignment Faking Mitigationsalignment.anthropic.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- “Alignment Faking” frame is somewhat fake — LessWronglesswrong.com
- [2412.14093] Alignment faking in large language modelsarxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- [2412.14093] Alignment faking in large language modelsarxiv.org
- Teaching Claude Whyalignment.anthropic.com
- Will alignment-faking Claude accept a deal to reveal its misalignment? — LessWronglesswrong.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Towards training-time mitigations for alignment faking in RL — LessWronglesswrong.com