[2603.01192] Grokking as a Phase Transition between Competing Basins: a Singular Learning Theory Approach
Abstract:Grokking, the abrupt transition from memorization to generalisation after extended training, suggests the presence of competing solution basins with distinct statistical properties. We study this phenomenon through the lens of Singular Learning Theory (SLT), a Bayesian framework that characterizes the geometry of the loss landscape via the local learning coefficient (LLC), a measure of the local degeneracy of the loss surface. SLT links lower-LLC basins to higher posterior mass concentration and lower expected generalisation error. Leveraging this theory, we interpret grokking in quadratic networks as a phase transition between competing near-zero-loss solution basins. Our contributions are two-fold: we derive closed-form expressions for the LLC in quadratic networks trained on modular arithmetic tasks, with the corresponding empirical verification; as well as empirical evidence demonstrating that LLC trajectories provide a reliable tool for tracking generalisation dynamics and interpreting phase transitions during training.
[2603.01192] A Basin-Selection Perspective on Grokking via Singular Learning Theory Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Statistics > Machine Learning arXiv:2603.01192 (stat) [Submitted on 1 Mar 2026 ( v1 ), last revised 7 May 2026 (this version, v3)] Title: A Basin-Selection Perspective on Grokking via Singular Learning Theory Authors: Ben Cullen , Sergio Estan-Ruiz , Riya Danait , Jiayi Li View a PDF of the paper titled A Basin-Selection Perspective on Grokking via Singular Learning Theo
related reading
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- Towards Understanding Grokking: An Effective Theory of Representation Learningarxiv.org
- Clare Lyle | What's grokking good for?clarelyle.com
- DSLT 0. Distilling Singular Learning Theory — LessWronglesswrong.com
- Investigating the learning coefficient of modular addition: hackathon project — LessWronglesswrong.com
- [2605.01172] A Theory of Generalization in Deep Learningarxiv.org
- Neural networks generalize because of this one weird trick — LessWronglesswrong.com
- Growth and Form in a Toy Model of Superposition — LessWronglesswrong.com
- Do Machine Learning Models Memorize or Generalize?pair.withgoogle.com
- Neural networks generalize because of this one weird trick — AI Alignment Forumalignmentforum.org
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org