Notes on the Origin of Implicit Regularization in SGD
I wanted to highlight an intriguing paper I presented at a journal club recently: Samuel L Smith, Benoit Dherin, David Barrett, Soham De (2021) On the Origin of Implicit Regularization in Stochastic Gradient DescentThere's actually a related paper that came out simultaneously, studying full-batch gradient descent instead of SGD: David
April 1, 2021 · deep learning generalization SGD differerntial equations Notes on the Origin of Implicit Regularization in SGD I wanted to highlight an intriguing paper I presented at a journal club recently: Samuel L Smith, Benoit Dherin, David Barrett, Soham De (2021) On the Origin of Implicit Regularization in Stochastic Gradient Descent There's actually a related paper that came out simultaneously, studying full-batch gradient descent instead of SGD: David G.T. Barrett, Benoit Dherin (2021) Implicit Gradient Regularization One of the most important insights in machine learning over
Explore this link on the map →related reading
- [2101.12176] On the Origin of Implicit Regularization in Stochastic Gradient Descentarxiv.org
- Stochastic gradient descent - Wikipediaen.m.wikipedia.org
- [1609.04836] On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minimaarxiv.org
- Why Momentum Really Worksdistill.pub
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- The Little Book of Deep Learningfleuret.org
- The Generalization Mystery: Sharp vs Flat Minimainference.vc
- [2605.01172] A Theory of Generalization in Deep Learningarxiv.org
- Why Does SGD Love Flat Minima? Marginally Better blogrishit-dagli.github.io
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- A minimizer Far, Far Away – Parameter-free Learning and Optimization Algorithmsparameterfree.com
- Just Ask for Generalization | Eric Jangevjang.com