DeepSpeed ZeRO++: A leap in speed for LLM and chat model training with 4X less communication - Microsoft Research
Large AI models are transforming the digital world. Generative language models like Turing-NLG, ChatGPT, and GPT-4, powered by large language models (LLMs), are incredibly versatile, capable of performing tasks like summarization, coding, and translation. Similarly, large multimodal generative models like DALL·E, Microsoft Designer, and Bing Image Creator can generate art, architecture, videos, and other digital assets, empowering content creators, architects, and engineers to explore new frontiers of creative productivity. However, training these large models requires considerable memory and computing resources across hundreds or even thousands of GPU devices. For instance, training the Megatron-Turing NLG 530B model utilized over 4,000 NVidia A100 GPUs. Efficiently leveraging these resources requires a complex system of optimizations to partition the models into pieces that fit into the memory of individual devices, and to efficiently parallelize the computing across these devices. A
DeepSpeed ZeRO++: A leap in speed for LLM and chat model training with 4X less communication - Microsoft Research Skip to main content Research Publications Code & data People Microsoft Research blog Artificial intelligence Audio & acoustics Computer vision Graphics & multimedia Human-computer interaction Human language technologies Search & information retrieval Data platforms and analytics Hardware & devices Programming languages & software engineering Quantum computing Security, privacy & cryptography Systems & networking Algorithms Mathematics Ecology & environment Economics Medical, healt
Explore this link on the map →saved by
related reading
- How To Scale Your Modeljax-ml.github.io
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Everything about Distributed Training and Efficient Finetuning | Sumanth's Personal Websitesumanthrh.com
- Composer2.pdfcursor.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- LLM Resourcesforrestbicker.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- 5D parallelism in LLM training - gdymind's Bloggdymind.com
- GenAI Handbookgenai-handbook.github.io
- Orbit - Ultra-efficient RL Pipelinespherelab.ai