flâneur — a map of the web's best reading

Clare Lyle | What's grokking good for?

clarelyle.com · 2,118 words · saved by 1 readers

The traditional school of machine learning theory states that as your model class becomes more complex, you should expect whatever function your training algorithm finds in that model class to fit your training dataset will be worse, on average, at generalizing to new data. One passes neatly between three regimes: from under-fitting to fitting to over-fitting. Machine learning textbooks in 2017, when I was in school, particularly loved the following visualization. The justification for this figure is based on decades of beautiful work in statistical learning theory. For a brief primer, you can refer to earlier posts on e.g. PAC-Bayesian generalization bounds or model complexity measures for modern machine learning. The mathematics underlying the relationship between the empirical and expected risk are deep. They also utterly fail to predict the behaviour of modern deep learning systems. At first glance, this received wisdom isn’t completely at odds with the modern era of “jUsT aDd MorE

Clare Lyle | What's grokking good for? What's grokking good for? - Clare Lyle Posted on June 22, 2025 What's grokking good for? Some intuition on the what, why, and how of delayed generalization A primer on delayed generalization The traditional school of machine learning theory states that as your model class becomes more complex, you should expect whatever function your training algorithm finds in that model class to fit your training dataset will be worse, on average, at generalizing to new data. One passes neatly between three regimes: from under-fitting to fitting to over-fitting. Machine

Explore this link on the map →

saved by

related reading