Reducing Activation Recomputation
arxiv.org · 8,193 words · saved by 1 readers
N/A
Reducing Activation Recomputation in Large Transformer Models Vijay Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro NVIDIA arXiv:2205.05198v1 [cs.LG] 10 May 2022 Abstract…
related reading
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- How To Scale Your Modeljax-ml.github.io
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- 1910.02054v3arxiv.org
- 5D parallelism in LLM training - gdymind's Bloggdymind.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Paradigms of Parallelism | Colossal-AIcolossalai.org
- Overleaf Examplearxiv.org
- 2305.19370arxiv.org
- [1909.08053] Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelismarxiv.org
- Pipeline-Parallelism: Distributed Training via Model Partitioningsiboehm.com