Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior — LessWrong
lesswrong.com · 6,400 words · saved by 1 readers
This is a link post for two papers that came out today: …
x Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior — LessWrong MATS Program AI Frontpage 2025 Top Fifty: 11 % 176 Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior by Sam Marks , Nevan Wichers , Daniel Tan , Aram Ebtekar , Jozdien , David Africa , Alex Mallen , Fabien Roger 8th Oct 2025 AI Alignment Forum 2 min read 37 176 Ω 77 This is a link post for two papers that came out today: Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time ( Tan et al. ) Ino
related reading
- How hard is it to inoculate against misalignment generalization? — LessWronglesswrong.com
- [2510.04340] Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timearxiv.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- [2606.12016] Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalizationarxiv.org
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org