✳flâneur — a map of the web's best reading
Paradigms of Parallelism | Colossal-AI
colossalai.org · 1,338 words · saved by 1 readers
Author: Shenggui Li, Siqi Mai
On this page Paradigms of Parallelism Author: Shenggui Li, Siqi Mai Introduction With the development of deep learning, there is an increasing demand for parallel training. This is because that model and datasets are getting larger and larger and training time becomes a nightmare if we stick to single-GPU training. In this section, we will provide a brief overview of existing methods to parallelize training. If you wish to add on to this post, you may create a discussion in the GitHub forum . Data Parallel Data parallel is the most common form of parallelism due to its simplicity. In data
Explore this link on the map →saved by
related reading
- Parallelism in Distributed Deep Learning · Better Tomorrow with Computer Scienceinsujang.github.io
- Pipeline-Parallelism: Distributed Training via Model Partitioningsiboehm.com
- 5D parallelism in LLM training - gdymind's Bloggdymind.com
- How To Scale Your Modeljax-ml.github.io
- ml-engineering/model-parallelism at master · stas00/ml-engineering · GitHubgithub.com
- irhum.github.io - Tensor Parallelism with jax.pjitirhum.github.io
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Visualizing 6D Mesh Parallelism · mainmain-horse.github.io
- How to Parallelize Deep Learning on GPUs Part 1/2: Data Parallelism - Tim Dettmerstimdettmers.com
- Training Deep Networks with Data Parallelism in Jaxmishalaskin.com
- Everything about Distributed Training and Efficient Finetuning | Sumanth's Personal Websitesumanthrh.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com