Neural networks generalize because of this one weird trick — LessWrong
Produced under the mentorship of Evan Hubinger as part of the SERI ML Alignment Theory Scholars Program - Winter 2022 Cohort A big thank you to all of the people who gave me feedback on this post: Edmund Lau, Dan Murfet, Alexander Gietelink Oldenziel, Lucius Bushnaq, Rob Krzyzanowski, Alexandre Variengen, Jiri Hoogland, and Russell Goyder. Statistical learning theory is lying to you: "overparametrized" models actually aren't overparametrized, and generalization is not just a question of broad basins. To first order, that's because loss basins actually aren't basins but valleys, and at the base of these valleys lie "rivers" of constant, minimum loss. The higher the dimension of these minimum sets, the lower the effective dimensionality of your model.[1] Generalization is a balance between expressivity (more effective parameters) and simplicity (fewer effective parameters). In particular, it is the singularities of these minimum-loss sets — points at which the tangent vanishes — that det
x Neural networks generalize because of this one weird trick — LessWrong Best of LessWrong 2023 Singular Learning Theory Interpretability (ML & AI) MATS Program AI Frontpage 215 Neural networks generalize because of this one weird trick by Jesse Hoogland 18th Jan 2023 AI Alignment Forum Linkpost for www.jessehoogland.com 18 min read 35 215 Ω 73 Produced under the mentorship of Evan Hubinger as part of the SERI ML Alignment Theory Scholars Program - Winter 2022 Cohort A big thank you to all of the people who gave me feedback on this post: Edmund Lau, Dan Murfet, Alexander Gietelink Oldenziel, L
Explore this link on the map →related reading
- Neural networks generalize because of this one weird trick — AI Alignment Forumalignmentforum.org
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- DSLT 0. Distilling Singular Learning Theory — LessWronglesswrong.com
- arxiv.org/pdf/1805.08522arxiv.org
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- DSLT 2. Why Neural Networks obey Occam's Razor — LessWronglesswrong.com
- DSLT 2. Why Neural Networks obey Occam's Razor — LessWronglesswrong.com
- [2605.01172] A Theory of Generalization in Deep Learningarxiv.org
- The Scaling Hypothesis · Gwern.netgwern.net
- [2503.02113] Deep Learning is Not So Mysterious or Differentarxiv.org
- On neural scaling and the quanta hypothesisericjmichaud.com