Why Does SGD Love Flat Minima? Marginally Better blog
This post is an intuitive look through the chronicles of Stochastic Gradient Descent (SGD), one of the most popular algorithms. I try to explain the reasoning behind some of the interesting aspects of SGD, like the stochastic gradient noise, the SGN covariance matrix, and SGD’s preference for flat minima. Recently, I’ve found it fascinating to explore and draw parallels with learning dynamics of algorithms based on gradient descent. Some ideas have particularly taken a bit of time for me to truly understand and with this, I hope to make it easier for others. Let’s take a quick look at Stochastic Gradient Descent (SGD). We consider the data samples 𝑥 = { 𝑥 𝑗 } 𝑗 = 1 𝑚 and similarly 𝑦 = { 𝑦 𝑗 } 𝑗 = 1 𝑚 , the model parameters 𝜃 , and some loss function 𝐿 ( 𝜃 , 𝑥 , 𝑦 ) . And following popular literature, for brevity, we denote the training loss as 𝐿 ( 𝜃 ) . We have SGD as follows: where 𝜂 is the step size, and 𝐿 ^ ( 𝜃 𝑡 , 𝑥 , 𝑦 ) is the stochastic estimate
Table of Contents ▼ But first… Modelling the Stochastic Gradient Noise (SGN) SGD can reach multiple minima What affects the SGN covariance matrix? SGD chooses a flat minima Concluding Thoughts Citation References Some things I learned from and some things I tried from a course at UofT. This article is a look through the chronicles of Stochastic Gradient Descent (SGD), one of the most popular algorithms. I try to explain the reasoning behind some of the interesting aspects of SGD, like the stochastic gradient noise, the SGN covariance matrix, and SGD’s preference for flat minima. Recently, I’ve
Explore this link on the map →related reading
- Stochastic gradient descent - Wikipediaen.m.wikipedia.org
- [2101.12176] On the Origin of Implicit Regularization in Stochastic Gradient Descentarxiv.org
- Thoughts on loss landscapes and why deep learning worksberen.io
- Notes on the Origin of Implicit Regularization in SGDinference.vc
- Thoughts on Loss Landscapes and why Deep Learning works — LessWronglesswrong.com
- Why Momentum Really Worksdistill.pub
- A minimizer Far, Far Away – Parameter-free Learning and Optimization Algorithmsparameterfree.com
- A Visual Explanation of Gradient Descent Methods (Momentum, AdaGrad, RMSProp, Adam) | Towards Data Sciencetowardsdatascience.com
- [2503.22478] Almost Bayesian: The Fractal Dynamics of Stochastic Gradient Descentarxiv.org
- The Generalization Mystery: Sharp vs Flat Minimainference.vc
- [1512.04202] Preconditioned Stochastic Gradient Descentarxiv.org
- The Little Book of Deep Learningfleuret.org