Why Momentum Really Works
distill.pub · 17,083 words · saved by 7 readers
We often think of optimization with momentum as a ball rolling down a hill. This isn't wrong, but there is much more to the story.
Why Momentum Really Works Distill About Prize Submit ⋆ \star ⋆ + + + − - − = = = α \alpha α λ \lambda λ β \beta β R R R α = \alpha= α = β = \beta= β = β = 0 \beta = 0 β = 0 β = 1 \beta=1 β = 1 α = 1 / λ i \alpha = 1/\lambda_i α = 1 / λ i m o d e l \text{model} model 0 p 1 0 p_1 0 p 1 0 p ¯ 1 0 \bar{p}_1 0 p ¯ 1 2 β 2\sqrt{\beta} 2 √ β λ i \lambda_i λ i λ i = 0 \lambda_i = 0 λ i = 0 α > 1 / λ i \alpha > 1/\lambda_i α > 1 / λ i max { ∣ σ 1 ∣ , ∣ σ 2 ∣ } > 1 \max\{|\sigma_1|,|\sigma_2|\} > 1 max { ∣ σ 1 ∣ , ∣ σ 2 ∣ } > 1 x i k − x i
saved by
related reading
- A Visual Explanation of Gradient Descent Methods (Momentum, AdaGrad, RMSProp, Adam) | Towards Data Sciencetowardsdatascience.com
- Deriving Muonjeremybernste.in
- Modular Manifolds - Thinking Machines Labthinkingmachines.ai
- Gradient Descent With Momentum from Scratch - MachineLearningMastery.commachinelearningmastery.com
- Ji-Ha's Blogjiha-kim.github.io
- Mathematical optimization - Wikipediaen.wikipedia.org
- Gradient descent - Wikipediaen.wikipedia.org
- Muon: An optimizer for hidden layers in neural networks | Keller Jordan blogkellerjordan.github.io
- Understanding the Neural Tangent Kernel – EigenTaleseigentales.com
- Stochastic gradient descent - Wikipediaen.m.wikipedia.org
- Scaling laws of optimization – Machine Learning Research Blogfrancisbach.com
- [1711.05101] Decoupled Weight Decay Regularizationarxiv.org