[2411.19870] DeMo: Decoupled Momentum Optimization
Abstract:Scaling neural network training increasingly depends on synchronous data-parallelism, yet full-precision gradient all-reduce imposes a severe communication bottleneck. We propose Decoupled Momentum Optimization (DeMo), a drop-in replacement for any momentum-based optimizers that significantly reduces the communication bandwidth while maintaining convergence. DeMo (i) decouples local momentum updates, (ii) applies a fast orthonormal transform (e.g., DCT) followed by top-k sparsification, and (iii) reuses the momentum buffer as error feedback via momentum subtraction. This design reduces per-step communication by up to two orders of magnitude with minimal computational overhead. Experiments on 300M and 1B-parameter DeMo language models show DeMo transmits up to 85x less data per GPU than AdamW-DDP while achieving comparable loss and accuracy. DeMo is topology-agnostic and enables training across multi-datacenter or Ethernet-based setups. Code is available at this https URL
# link_25t4tnvm9on.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Bowen Peng; Lizhang Chen; Baiyu Su; Jeffrey Quesnelle; Diederik P. Kingma; Qiang Liu - Creator=arXiv GenPDF (tex2pdf:57610bf) - Custom.DOI=https://doi.org/10.48550/arXiv.2411.19870 - Custom.License=http://creativecommons.org/licenses/by/4.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.28 (TeX Live 2025) kpathsea version 6.4.1 - Custom.arXivID=https://arxiv.org/abs/2411.19870v2 - P
Explore this link on the map →saved by
related reading
- How To Scale Your Modeljax-ml.github.io
- Deriving Muonjeremybernste.in
- 5D parallelism in LLM training - gdymind's Bloggdymind.com
- Making Deep Learning go Brrrr From First Principleshorace.io
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Modular Manifolds - Thinking Machines Labthinkingmachines.ai
- The Little Book of Deep Learningfleuret.org
- Online KL Shampoo | Tildeblog.tilderesearch.com
- Why Momentum Really Worksdistill.pub
- Pipeline-Parallelism: Distributed Training via Model Partitioningsiboehm.com
- Overleaf Examplearxiv.org