1910.02054v3
arxiv.org · 8,777 words · saved by 1 readers
N/A
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models Samyam Rajbhandari∗ , Jeff Rasley∗, Olatunji Ruwase, Yuxiong He arXiv:1910.02054v3 [cs.LG] 13 May 2020 {samyamr, jerasley, olruwase, yuxhe}@microsoft.com Abstract Large deep learning models offer significant accuracy gains, but training billions to trillions…
saved by
related reading
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- How To Scale Your Modeljax-ml.github.io
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- DeepSpeed ZeRO++: A leap in speed for LLM and chat model training with 4X less communication - Microsoft Researchmicrosoft.com
- 5D parallelism in LLM training - gdymind's Bloggdymind.com
- Everything about Distributed Training and Efficient Finetuning | Sumanth's Personal Websitesumanthrh.com
- Paradigms of Parallelism | Colossal-AIcolossalai.org
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- [1909.08053] Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelismarxiv.org
- The Little Book of Deep Learningfleuret.org
- Pipeline-Parallelism: Distributed Training via Model Partitioningsiboehm.com
- Reducing Activation Recomputationarxiv.org