Generalisation Hacking: a first look at adversarial generalisation failures in deliberative alignment — LessWrong
Background Deliberative alignment is a powerful post-training alignment technique that involves generating and training on re-contextualised supervis…
x Generalisation Hacking: a first look at adversarial generalisation failures in deliberative alignment — LessWrong AI Frontpage 54 Generalisation Hacking: a first look at adversarial generalisation failures in deliberative alignment by Cam , Puria 17th Nov 2025 AI Alignment Forum 10 min read 2 54 Ω 16 Background Deliberative alignment is a powerful post-training alignment technique that involves generating and training on re-contextualised supervised fine-tuning (SFT) datasets generated with a set of principles in context. The process takes three steps: With the set of principles [1] (hencefo
Explore this link on the map →related reading
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- How hard is it to inoculate against misalignment generalization? — LessWronglesswrong.com
- Alignment faking in large language modelsarxiv.org
- Natural emergent misalignment from reward hacking \ Anthropicanthropic.com
- How far does alignment midtraining generalize?alignment.openai.com