flâneur — a map of the web's best reading

Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior — LessWrong

lesswrong.com · 6,400 words · saved by 1 readers

This is a link post for two papers that came out today: …

x Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior — LessWrong MATS Program AI Frontpage 2025 Top Fifty: 11 % 176 Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior by Sam Marks , Nevan Wichers , Daniel Tan , Aram Ebtekar , Jozdien , David Africa , Alex Mallen , Fabien Roger 8th Oct 2025 AI Alignment Forum 2 min read 37 176 Ω 77 This is a link post for two papers that came out today: Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time ( Tan et al. ) Ino

Explore this link on the map →

related reading