✳flâneur — a map of the web's best reading
Zipfian grokking | Jasper Gilley
jagilley.github.io · 2,626 words · saved by 2 readers
A toy problem where wide abstractions fight deep ones — and a test bed for measuring the gap between dataset MDL and data-generating process MDL.
Zipfian grokking | Jasper Gilley ← All Posts Zipfian grokking Jasper Gilley TLDR: we modify the data distribution of a typical modular addition grokking setup to follow Zipf's Law, and show that it leads to persistent, predictable instability in grokked models' behavior. The imbalanced distribution causes the model to oscillate between a more generalizing and a more memorizing solution. We believe that this is a promising toy problem for testing new methods of eliciting generalization. One of the reasons for the successes of the scaling era of AI has been passive regularization: throwing
Explore this link on the map →saved by
related reading
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- [2605.01172] A Theory of Generalization in Deep Learningarxiv.org
- NL.pdfabehrouz.github.io
- On neural scaling and the quanta hypothesisericjmichaud.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Clare Lyle | What's grokking good for?clarelyle.com
- Do Machine Learning Models Memorize or Generalize?pair.withgoogle.com
- The Scaling Hypothesis · Gwern.netgwern.net
- The Little Book of Deep Learningfleuret.org
- Just Ask for Generalization | Eric Jangevjang.com
- A Mechanistic Interpretability Analysis of Grokking — LessWronglesswrong.com