✳flâneur — a map of the web's best reading
Why Momentum Really Works
distill.pub · 17,083 words · saved by 5 readers
We often think of optimization with momentum as a ball rolling down a hill. This isn't wrong, but there is much more to the story.
Why Momentum Really Works Distill About Prize Submit ⋆ \star ⋆ + + + − - − = = = α \alpha α λ \lambda λ β \beta β R R R α = \alpha= α = β = \beta= β = β = 0 \beta = 0 β = 0 β = 1 \beta=1 β = 1 α = 1 / λ i \alpha = 1/\lambda_i α = 1 / λ i m o d e l \text{model} model 0 p 1 0 p_1 0 p 1 0 p ¯ 1 0 \bar{p}_1 0 p ¯ 1 2 β 2\sqrt{\beta} 2 √ β λ i \lambda_i λ i λ i = 0 \lambda_i = 0 λ i = 0 α > 1 / λ i \alpha > 1/\lambda_i α > 1 / λ i max { ∣ σ 1 ∣ , ∣ σ 2 ∣ } > 1 \max\{|\sigma_1|,|\sigma_2|\} > 1 max { ∣ σ 1 ∣ , ∣ σ 2 ∣ } > 1 x i k − x i
Explore this link on the map →saved by
related reading
- A Visual Explanation of Gradient Descent Methods (Momentum, AdaGrad, RMSProp, Adam) | Towards Data Sciencetowardsdatascience.com
- Modular Manifolds - Thinking Machines Labthinkingmachines.ai
- Deriving Muonjeremybernste.in
- Gradient Descent With Momentum from Scratch - MachineLearningMastery.commachinelearningmastery.com
- Understanding the Neural Tangent Kernel – EigenTaleseigentales.com
- Gradient descent - Wikipediaen.wikipedia.org
- Stochastic gradient descent - Wikipediaen.m.wikipedia.org
- Scaling laws of optimization – Machine Learning Research Blogfrancisbach.com
- An overview of gradient descent optimization algorithmsruder.io
- Effortless optimization through gradient flows – Machine Learning Research Blogfrancisbach.com
- Understanding Deep Learning Optimizers: Momentum, AdaGrad, RMSProp & Adam | Towards Data Sciencetowardsdatascience.com
- On the Link Between Polynomials and Optimization, Part 1fa.bianp.net