Generalisation Hacking: a first look at adversarial generalisation failures in deliberative alignment — LessWrong
Background Deliberative alignment is a powerful post-training alignment technique that involves generating and training on re-contextualised supervis…
x Generalisation Hacking: a first look at adversarial generalisation failures in deliberative alignment — LessWrong AI Frontpage 54 Generalisation Hacking: a first look at adversarial generalisation failures in deliberative alignment by Cam , Puria 17th Nov 2025 AI Alignment Forum 10 min read 2 54 Ω 16 Background Deliberative alignment is a powerful post-training alignment technique that involves generating and training on re-contextualised supervised fine-tuning (SFT) datasets generated with a set of principles in context. The process takes three steps: With the set of principles [1] (hencefo
related reading
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- [2606.12016] Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalizationarxiv.org
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- How hard is it to inoculate against misalignment generalization? — LessWronglesswrong.com
- Reinforcement learning towards broadly and persistently beneficial modelsalignment.openai.com
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com