[2603.01192] Grokking as a Phase Transition between Competing Basins: a Singular Learning Theory Approach
Abstract:Grokking, the abrupt transition from memorization to generalisation after extended training, suggests the presence of competing solution basins with distinct statistical properties. We study this phenomenon through the lens of Singular Learning Theory (SLT), a Bayesian framework that characterizes the geometry of the loss landscape via the local learning coefficient (LLC), a measure of the local degeneracy of the loss surface. SLT links lower-LLC basins to higher posterior mass concentration and lower expected generalisation error. Leveraging this theory, we interpret grokking in quadratic networks as a phase transition between competing near-zero-loss solution basins. Our contributions are two-fold: we derive closed-form expressions for the LLC in quadratic networks trained on modular arithmetic tasks, with the corresponding empirical verification; as well as empirical evidence demonstrating that LLC trajectories provide a reliable tool for tracking generalisation dynamics and interpreting phase transitions during training.
[2603.01192] A Basin-Selection Perspective on Grokking via Singular Learning Theory Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Statistics > Machine Learning arXiv:2603.01192 (stat) [Submitted on 1 Mar 2026 ( v1 ), last revised 7 May 2026 (this version, v3)] Title: A Basin-Selection Perspective on Grokking via Singular Learning Theory Authors: Ben Cullen , Sergio Estan-Ruiz , Riya Danait , Jiayi Li View a PDF of the paper titled A Basin-Selection Perspective on Grokking via Singular Learning Theo
Explore this link on the map →related reading
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- Clare Lyle | What's grokking good for?clarelyle.com
- DSLT 0. Distilling Singular Learning Theory — LessWronglesswrong.com
- [2605.01172] A Theory of Generalization in Deep Learningarxiv.org
- Growth and Form in a Toy Model of Superposition — LessWronglesswrong.com
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- Neural networks generalize because of this one weird trick — LessWronglesswrong.com
- Investigating the learning coefficient of modular addition: hackathon project — LessWronglesswrong.com
- Neural networks generalize because of this one weird trick — AI Alignment Forumalignmentforum.org
- Zipfian grokking | Jasper Gilleyjagilley.github.io
- Elon Litman | Elements of a Vector Spaceelonlit.com