Hypothesis: gradient descent prefers general circuits - LessWrong
Summary: I discuss a potential mechanistic explanation for why SGD might prefer general circuits for generating model outputs. I use this preference to explain how models can learn to generalize even…
x Hypothesis: gradient descent prefers general circuits — LessWrong Gradient Descent Optimization AI Frontpage 46 Hypothesis: gradient descent prefers general circuits by Quintin Pope 8th Feb 2022 AI Alignment Forum 14 min read 26 46 Ω 20 Summary: I discuss a potential mechanistic explanation for why SGD might prefer general circuits for generating model outputs. I use this preference to explain how models can learn to generalize even after overfitting to near zero training error (i.e., grokking). I also discuss other perspectives on grokking and deep learning generalization. Additionally, I d
Explore this link on the map →related reading
- Clare Lyle | What's grokking good for?clarelyle.com
- The Scaling Hypothesis · Gwern.netgwern.net
- Just Ask for Generalization | Eric Jangevjang.com
- Do Machine Learning Models Memorize or Generalize?pair.withgoogle.com
- The Little Book of Deep Learningfleuret.org
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- Thoughts on Loss Landscapes and why Deep Learning works — LessWronglesswrong.com
- Neural networks and deep learningneuralnetworksanddeeplearning.com
- arxiv.org/pdf/1805.08522arxiv.org
- [2605.01172] A Theory of Generalization in Deep Learningarxiv.org
- Zipfian grokking | Jasper Gilleyjagilley.github.io