Zipfian grokking | Jasper Gilley
jagilley.github.io · 2,626 words · saved by 2 readers
A toy problem where wide abstractions fight deep ones — and a test bed for measuring the gap between dataset MDL and data-generating process MDL.
Zipfian grokking | Jasper Gilley ← All Posts Zipfian grokking Jasper Gilley TLDR: we modify the data distribution of a typical modular addition grokking setup to follow Zipf's Law, and show that it leads to persistent, predictable instability in grokked models' behavior. The imbalanced distribution causes the model to oscillate between a more generalizing and a more memorizing solution. We believe that this is a promising toy problem for testing new methods of eliciting generalization. One of the reasons for the successes of the scaling era of AI has been passive regularization: throwing
saved by
related reading
- [2608.13335] Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Lawsarxiv.org
- Do Machine Learning Models Memorize or Generalize?pair.withgoogle.com
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- [2605.01172] A Theory of Generalization in Deep Learningarxiv.org
- NL.pdfabehrouz.github.io
- On neural scaling and the quanta hypothesisericjmichaud.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- Clare Lyle | What's grokking good for?clarelyle.com
- The Scaling Hypothesis · Gwern.netgwern.net
- Towards Understanding Grokking: An Effective Theory of Representation Learningarxiv.org