✳flâneur — a map of the web's best reading
How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? — LessWrong
lesswrong.com · 7,148 words · saved by 1 readers
Authors: Dylan Xu, Alek Westover, Vivek Hebbar, Sebastian Prasanna, Nathan Sheffield, Buck Shlegeris, Julian Stastny …
x How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? — LessWrong AI Frontpage 62 How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? by Dylan Xu , Alek Westover , Vivek Hebbar , SebastianP , frisby , Julian Stastny 20th Apr 2026 24 min read 5 62 Authors: Dylan Xu, Alek Westover, Vivek Hebbar, Sebastian Prasanna, Nathan Sheffield, Buck Shlegeris, Julian Stastny Thanks to Eric Gan and Aghyad Deeb for feedback on a draft of this post. EDIT: After further consideration and @nostal
Explore this link on the map →saved by
related reading
- The persona selection model — LessWronglesswrong.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- Model organisms researchers should check whether high LRs defeat their model organisms — LessWronglesswrong.com
- [2602.05910] Chunky Post-Training: Data Driven Failures of Generalizationarxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- [2510.04340] Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timearxiv.org
- [2506.19733] Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?arxiv.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai
- Generalization Dynamics of LM Pre-training — Jiaxin Wenjiaxin-wen.github.io
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org