✳flâneur — a map of the web's best reading
How hard is it to inoculate against misalignment generalization? — LessWrong
lesswrong.com · 4,750 words · saved by 2 readers
TL;DR: Simple inoculation prompts that prevent misalignment generalization in toy setups don't scale to more realistic reward hacking. When I fine-tu…
x How hard is it to inoculate against misalignment generalization? — LessWrong Deceptive Alignment Language Models (LLMs) AI Frontpage 46 How hard is it to inoculate against misalignment generalization? by Jozdien 6th Jan 2026 AI Alignment Forum 16 min read 4 46 Ω 22 TL;DR: Simple inoculation prompts that prevent misalignment generalization in toy setups don't scale to more realistic reward hacking. When I fine-tuned models on realistic reward hacks , only prompts close to the dataset generation prompt were sufficient to prevent misalignment. This seems like a specification problem: the model
Explore this link on the map →saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior — LessWronglesswrong.com
- Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Natural emergent misalignment from reward hacking \ Anthropicanthropic.com
- [2510.04340] Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timearxiv.org
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Paper: Prompt Optimization Makes Misalignment Legible — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com