Shampoo: Preconditioned Stochastic Tensor Optimization | HTML5
Preconditioned gradient methods are among the most general and powerful tools in optimization. However, preconditioning requires storing and manipulating prohibitively large matrices. We describe and analyze a new structure-aware preconditioning algorithm, called Shampoo, for stochastic optimization over tensor spaces. Shampoo maintains a set of preconditioning matrices, each of which operates on a single dimension, contracting over the remaining dimensions. We establish convergence guarantees in the stochastic convex setting, the proof of which builds upon matrix trace inequalities. Our experiments with state-of-the-art deep learning models show that Shampoo is capable of converging considerably faster than commonly used optimizers. Although it involves a more complex update rule, Shampoo’s runtime per step is comparable to that of simple gradient methods such as SGD, AdaGrad, and Adam. Over the last decade, stochastic first-order optimization methods have emerged as the canonical too
Shampoo: Preconditioned Stochastic Tensor Optimization Vineet Gupta Tomer Koren † † footnotemark: Yoram Singer Google Brain. Email: {vineet,tkoren}@google.com Princeton University and Google Brain. Email: y.s@cs.princeton.edu Abstract Preconditioned gradient methods are among the most general and powerful tools in optimization. However, preconditioning requires storing and manipulating prohibitively large matrices. We describe and analyze a new structure-aware preconditioning algorithm, called Shampoo, for stochastic optimization over tensor spaces. Shampoo maintains a set of precondit
Explore this link on the map →related reading
- Online KL Shampoo | Tildeblog.tilderesearch.com
- [1512.04202] Preconditioned Stochastic Gradient Descentarxiv.org
- Modular Manifolds - Thinking Machines Labthinkingmachines.ai
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Why Momentum Really Worksdistill.pub
- Deriving Muonjeremybernste.in
- Preconditioner - Wikipediaen.wikipedia.org
- Your Transformer is Secretly an EOT Solver | Elements of a Vector Spaceelonlit.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Stochastic gradient descent - Wikipediaen.m.wikipedia.org
- 2404.17625arxiv.org
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io