✳flâneur — a map of the web's best reading
Ambiguous out-of-distribution generalization on an algorithmic task — LessWrong
lesswrong.com · 4,790 words · saved by 1 readers
Introduction It's now well known that simple neural network models often "grok" algorithmic tasks. That is, when trained for many epochs on a subset…
x Ambiguous out-of-distribution generalization on an algorithmic task — LessWrong Deceptive Alignment Distributional Shifts Grokking (ML) Interpretability (ML & AI) MATS Program AI Frontpage 84 Ambiguous out-of-distribution generalization on an algorithmic task by Wilson Wu , Louis Jaburi 13th Feb 2025 13 min read 6 84 Introduction It's now well known that simple neural network models often "grok" algorithmic tasks. That is, when trained for many epochs on a subset of the full input space, the model quickly attains perfect train accuracy and then, much later, near-perfect test accuracy. In the
Explore this link on the map →related reading
- Clare Lyle | What's grokking good for?clarelyle.com
- Just Ask for Generalization | Eric Jangevjang.com
- Do Machine Learning Models Memorize or Generalize?pair.withgoogle.com
- Zipfian grokking | Jasper Gilleyjagilley.github.io
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- arxiv.org/pdf/1805.08522arxiv.org
- The Little Book of Deep Learningfleuret.org
- AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWronglesswrong.com
- [2605.01172] A Theory of Generalization in Deep Learningarxiv.org
- A Mechanistic Interpretability Analysis of Grokking — LessWronglesswrong.com
- Generalization Dynamics of LM Pre-training — Jiaxin Wenjiaxin-wen.github.io
- [1911.08731] Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalizationarxiv.org