Thoughts on loss landscapes and why deep learning works
Epistemic status: Pretty uncertain. I don’t have an expert level understanding of current views in the science of deep learning about why optimization works but just read papers as an amateur. Some of the arguments I present here might be already either known or disproven. If so please let me know! There are essentially two fundamental questions in the science of deep learning: 1.) Why are models trainable? And 2.) Why do models generalize? The answer to both of these questions relates to the nature and basic geometry of the loss landscape which is ultimately determined by the computational architecture of the model. Here I present my personal and fairly idiosyncratic and speculative answers to these questions and present what I think are fairly novel answers for both of these questions. Let’s get started. First let’s think about the first question: Why are deep learning models trainable at all? A-priori, they shouldn’t be. Deep neural networks are fundamentally solving extremely high
Epistemic status : Pretty uncertain. I don’t have an expert level understanding of current views in the science of deep learning about why optimization works but just read papers as an amateur. Some of the arguments I present here might be already either known or disproven. If so please let me know! There are essentially two fundamental questions in the science of deep learning: 1.) Why are models trainable? And 2.) Why do models generalize? The answer to both of these questions relates to the nature and basic geometry of the loss landscape which is ultimately determined by the computational a
related reading
- Thoughts on Loss Landscapes and why Deep Learning works — LessWronglesswrong.com
- The Little Book of Deep Learningfleuret.org
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- Statistical Mechanics of Deep Learningganguli-gang.stanford.edu
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- Why Deep Learning Works – Key Insights and Saddle Points - KDnuggetskdnuggets.com
- The Decade of Deep Learning | Leo Gaobmk.sh
- The Generalization Mystery: Sharp vs Flat Minimainference.vc
- 1412.0233arxiv.org
- Loss Landscape | A.I deep learning explorations of morphology & dynamicslosslandscape.com