How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? — LessWrong
lesswrong.com · 7,148 words · saved by 1 readers
Authors: Dylan Xu, Alek Westover, Vivek Hebbar, Sebastian Prasanna, Nathan Sheffield, Buck Shlegeris, Julian Stastny …
x How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? — LessWrong AI Frontpage 62 How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? by Dylan Xu , Alek Westover , Vivek Hebbar , SebastianP , frisby , Julian Stastny 20th Apr 2026 24 min read 5 62 Authors: Dylan Xu, Alek Westover, Vivek Hebbar, Sebastian Prasanna, Nathan Sheffield, Buck Shlegeris, Julian Stastny Thanks to Eric Gan and Aghyad Deeb for feedback on a draft of this post. EDIT: After further consideration and @nostal
saved by
related reading
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Model organisms researchers should check whether high LRs defeat their model organisms — LessWronglesswrong.com
- [2602.05910] Chunky Post-Training: Data Driven Failures of Generalizationarxiv.org
- [2606.12016] Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalizationarxiv.org
- [2510.04340] Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timearxiv.org
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimesarxiv.org
- [2506.19733] Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?arxiv.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Modelsarxiv.org
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai