✳flâneur — a map of the web's best reading
Do Machine Learning Models Memorize or Generalize?
pair.withgoogle.com · 5,261 words · saved by 1 readers
An interactive introduction to grokking and mechanistic interpretability.
Do Machine Learning Models Memorize or Generalize? Explorables Do Machine Learning Models Memorize or Generalize? By Adam Pearce , Asma Ghandeharioun , Nada Hussein , Nithum Thain , Martin Wattenberg and Lucas Dixon August 2023 In 2021, researchers made a striking discovery while training a series of tiny models on toy tasks . They found a set of models that suddenly flipped from memorizing their training data to correctly generalizing on unseen inputs after training for much longer. This phenomenon – where generalization seems to happen abruptly and long after fitting the training data – is c
Explore this link on the map →related reading
- Clare Lyle | What's grokking good for?clarelyle.com
- Zipfian grokking | Jasper Gilleyjagilley.github.io
- A Mechanistic Interpretability Analysis of Grokking — AI Alignment Forumalignmentforum.org
- A Mechanistic Interpretability Analysis of Grokking — LessWronglesswrong.com
- Understanding Memorization via Loss Curvaturegoodfire.ai
- Just Ask for Generalization | Eric Jangevjang.com
- arxiv.org/pdf/2505.24832arxiv.org
- The Little Book of Deep Learningfleuret.org
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- NL.pdfabehrouz.github.io
- Hypothesis: gradient descent prefers general circuits — LessWronglesswrong.com
- Ambiguous out-of-distribution generalization on an algorithmic task — LessWronglesswrong.com