Thoughts on loss landscapes and why deep learning works
Epistemic status: Pretty uncertain. I don’t have an expert level understanding of current views in the science of deep learning about why optimization works but just read papers as an amateur. Some of the arguments I present here might be already either known or disproven. If so please let me know! There are essentially two fundamental questions in the science of deep learning: 1.) Why are models trainable? And 2.) Why do models generalize? The answer to both of these questions relates to the nature and basic geometry of the loss landscape which is ultimately determined by the computational architecture of the model. Here I present my personal and fairly idiosyncratic and speculative answers to these questions and present what I think are fairly novel answers for both of these questions. Let’s get started. First let’s think about the first question: Why are deep learning models trainable at all? A-priori, they shouldn’t be. Deep neural networks are fundamentally solving extremely high
Epistemic status : Pretty uncertain. I don’t have an expert level understanding of current views in the science of deep learning about why optimization works but just read papers as an amateur. Some of the arguments I present here might be already either known or disproven. If so please let me know! There are essentially two fundamental questions in the science of deep learning: 1.) Why are models trainable? And 2.) Why do models generalize? The answer to both of these questions relates to the nature and basic geometry of the loss landscape which is ultimately determined by the computational a
Explore this link on the map →related reading
- Thoughts on Loss Landscapes and why Deep Learning works — LessWronglesswrong.com
- The Little Book of Deep Learningfleuret.org
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- Statistical Mechanics of Deep Learningganguli-gang.stanford.edu
- The Decade of Deep Learning | Leo Gaobmk.sh
- Why Deep Learning Works – Key Insights and Saddle Points - KDnuggetskdnuggets.com
- Loss Landscape | A.I deep learning explorations of morphology & dynamicslosslandscape.com
- [2605.01172] A Theory of Generalization in Deep Learningarxiv.org
- Neural networks and deep learningneuralnetworksanddeeplearning.com
- arxiv.org/pdf/1805.08522arxiv.org