✳flâneur — a map of the web's best reading
[1511.06251] Stochastic modified equations and adaptive stochastic gradient algorithms
arxiv.org · saved by 1 readers
We develop the method of stochastic modified equations (SME), in which stochastic gradient algorithms are approximated in the weak sense by continuous-time stochastic differential equations. We exploit the continuous formulation together with optimal control theory to derive novel adaptive hyper-parameter adjustment policies. Our algorithms have competitive performance with the added benefit of being robust to varying models and datasets. This provides a general methodology for the analysis and design of stochastic gradient algorithms.
Explore this link on the map →related reading
- Stochastic gradient descent - Wikipediaen.m.wikipedia.org
- Policy Gradient Algorithms | Lil'Loglilianweng.github.io
- AdaGrad - Cornell University Computational Optimization Open Textbook - Optimization Wikioptimization.cbe.cornell.edu
- [1512.04202] Preconditioned Stochastic Gradient Descentarxiv.org
- A Review of Modern Stochastic Modeling: SDE/SPDE Numerics, Data-Driven Identification, and Generative Methods with Applications in Biology and Epidemiologyarxiv.org
- An SDE Framework for Adversarial Training, with Convergence and Robustness Analysisarxiv.org
- Highly optimized optimizers - by Ben Recht - arg minargmin.net
- Why Momentum Really Worksdistill.pub
- A Visual Explanation of Gradient Descent Methods (Momentum, AdaGrad, RMSProp, Adam) | Towards Data Sciencetowardsdatascience.com
- [2101.12176] On the Origin of Implicit Regularization in Stochastic Gradient Descentarxiv.org
- Part 3: Intro to Policy Optimization - Spinning Up documentationspinningup.openai.com
- A Primer on Stochastic Partial Differential Equationsmath.utah.edu