flâneur — a map of the web's best reading

How hard is it to inoculate against misalignment generalization? — LessWrong

lesswrong.com · 4,750 words · saved by 2 readers

TL;DR: Simple inoculation prompts that prevent misalignment generalization in toy setups don't scale to more realistic reward hacking. When I fine-tu…

x How hard is it to inoculate against misalignment generalization? — LessWrong Deceptive Alignment Language Models (LLMs) AI Frontpage 46 How hard is it to inoculate against misalignment generalization? by Jozdien 6th Jan 2026 AI Alignment Forum 16 min read 4 46 Ω 22 TL;DR: Simple inoculation prompts that prevent misalignment generalization in toy setups don't scale to more realistic reward hacking. When I fine-tuned models on realistic reward hacks , only prompts close to the dataset generation prompt were sufficient to prevent misalignment. This seems like a specification problem: the model

Explore this link on the map →

saved by

related reading