flâneur — a map of the web's best reading

Generalization Dynamics of LM Pre-training — Jiaxin Wen

jiaxin-wen.github.io · 4,894 words · saved by 1 readers

People typically assume that LMs stably mature from pattern-matching parrots to generalizable intelligence during pre-training. We build a toy eval suite and show this mental model is wrong: throughout pre-training, LMs frequently and suddenly hop between parrot-like and intelligence-like modes, i.e. distinct algorithms implemented by distinct circuits. We call this mode-hopping. Across our suite, LMs can suddenly latch onto memorized or in-context patterns instead of in-context learning, use System 1 instead of System 2 thinking, pick up what sounds true instead of what is true, fail at multi-hop persona QA, out-of-context reasoning, and emergent misalignment — then just as suddenly revert and generalize. Mode-hopping is not explained by standard optimization dynamics: it is locally stable and can not be fixed by checkpoint averaging. We instead think of it as a capacity allocation problem: in a capacity-bounded model, generalizable circuits must compete with the shallow ones learned

Generalization Dynamics of LM Pre-training — Jiaxin Wen ← back Generalization Dynamics of LM Pre-training Jiaxin Wen 1 , Zhengxuan Wu 2 , Dawn Song 1 , Lijie Chen 1 1 UC Berkeley · 2 Stanford (now at Google DeepMind) May 2026 Abstract People typically assume that LMs stably mature from pattern-matching parrots to generalizable intelligence during pre-training. We build a toy eval suite and show this mental model is wrong: throughout pre-training, LMs frequently and suddenly hop between parrot-like and intelligence-like modes, i.e. distinct algorithms implemented by distinct circuits. We call t

Explore this link on the map →

saved by

related reading