✳flâneur — a map of the web's best reading
Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior — LessWrong
lesswrong.com · 6,400 words · saved by 1 readers
This is a link post for two papers that came out today: …
x Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior — LessWrong MATS Program AI Frontpage 2025 Top Fifty: 11 % 176 Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior by Sam Marks , Nevan Wichers , Daniel Tan , Aram Ebtekar , Jozdien , David Africa , Alex Mallen , Fabien Roger 8th Oct 2025 AI Alignment Forum 2 min read 37 176 Ω 77 This is a link post for two papers that came out today: Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time ( Tan et al. ) Ino
Explore this link on the map →related reading
- How hard is it to inoculate against misalignment generalization? — LessWronglesswrong.com
- [2510.04340] Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timearxiv.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs — LessWronglesswrong.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Paper: Prompt Optimization Makes Misalignment Legible — LessWronglesswrong.com
- How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org