Generalization Dynamics of LM Pre-training — Jiaxin Wen
People typically assume that LMs stably mature from pattern-matching parrots to generalizable intelligence during pre-training. We build a toy eval suite and show this mental model is wrong: throughout pre-training, LMs frequently and suddenly hop between parrot-like and intelligence-like modes, i.e. distinct algorithms implemented by distinct circuits. We call this mode-hopping. Across our suite, LMs can suddenly latch onto memorized or in-context patterns instead of in-context learning, use System 1 instead of System 2 thinking, pick up what sounds true instead of what is true, fail at multi-hop persona QA, out-of-context reasoning, and emergent misalignment — then just as suddenly revert and generalize. Mode-hopping is not explained by standard optimization dynamics: it is locally stable and can not be fixed by checkpoint averaging. We instead think of it as a capacity allocation problem: in a capacity-bounded model, generalizable circuits must compete with the shallow ones learned
Generalization Dynamics of LM Pre-training — Jiaxin Wen ← back Generalization Dynamics of LM Pre-training Jiaxin Wen 1 , Zhengxuan Wu 2 , Dawn Song 1 , Lijie Chen 1 1 UC Berkeley · 2 Stanford (now at Google DeepMind) May 2026 Abstract People typically assume that LMs stably mature from pattern-matching parrots to generalizable intelligence during pre-training. We build a toy eval suite and show this mental model is wrong: throughout pre-training, LMs frequently and suddenly hop between parrot-like and intelligence-like modes, i.e. distinct algorithms implemented by distinct circuits. We call t
Explore this link on the map →saved by
related reading
- How far does alignment midtraining generalize?alignment.openai.com
- Just Ask for Generalization | Eric Jangevjang.com
- [2602.05910] Chunky Post-Training: Data Driven Failures of Generalizationarxiv.org
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? — LessWronglesswrong.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- [2502.19402] General Reasoning Requires Learning to Reason from the Get-goar5iv.labs.arxiv.org
- [2506.19733] Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?arxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai