DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
arxiv.org · 4,854 words · saved by 1 readers
N/A
D EEP S PEED U LYSSES : S YSTEM O PTIMIZATIONS FOR E NABLING T RAINING OF E XTREME L ONG S EQUENCE T RANSFORMER M ODELS Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang, Minjia Zhang arXiv:2309.14509v2 [cs.LG] 4 Oct 2023 Shuaiwen Leon Song, Samyam Rajbhandari, Yuxiong He…
related reading
- How To Scale Your Modeljax-ml.github.io
- LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelismarxiv.org
- Paradigms of Parallelism | Colossal-AIcolossalai.org
- 2305.19370arxiv.org
- 5D parallelism in LLM training - gdymind's Bloggdymind.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- [2205.14135] FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awarenessarxiv.org
- 1910.02054v3arxiv.org
- [2011.04006] Long Range Arena: A Benchmark for Efficient Transformersarxiv.org
- [1909.08053] Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelismarxiv.org
- Overleaf Examplearxiv.org