flâneur — a map of the web's best reading

Zipfian grokking | Jasper Gilley

jagilley.github.io · 2,626 words · saved by 2 readers

A toy problem where wide abstractions fight deep ones — and a test bed for measuring the gap between dataset MDL and data-generating process MDL.

Zipfian grokking | Jasper Gilley ← All Posts Zipfian grokking Jasper Gilley TLDR: we modify the data distribution of a typical modular addition grokking setup to follow Zipf's Law, and show that it leads to persistent, predictable instability in grokked models' behavior. The imbalanced distribution causes the model to oscillate between a more generalizing and a more memorizing solution. We believe that this is a promising toy problem for testing new methods of eliciting generalization. One of the reasons for the successes of the scaling era of AI has been passive regularization: throwing

Explore this link on the map →

saved by

related reading