✳flâneur — a map of the web's best reading
One weird trick for parallelizing convolutional neural networks | PDF
arxiv.org · 3,960 words · saved by 1 readers
N/A
One weird trick for parallelizing convolutional neural networks Alex Krizhevsky Google Inc. akrizhevsky@google.com April 29, 2014 arXiv:1404.5997v2 [cs.NE] 26 Apr 2014…
Explore this link on the map →related reading
- How to Parallelize Deep Learning on GPUs Part 1/2: Data Parallelism - Tim Dettmerstimdettmers.com
- Pipeline-Parallelism: Distributed Training via Model Partitioningsiboehm.com
- CS231n Deep Learning for Computer Visioncs231n.github.io
- Paradigms of Parallelism | Colossal-AIcolossalai.org
- Parallelism in Distributed Deep Learning · Better Tomorrow with Computer Scienceinsujang.github.io
- How To Scale Your Modeljax-ml.github.io
- [1706.02677] Accurate, Large Minibatch SGD: Training ImageNet in 1 Hourarxiv-vanity.com
- How to Parallelize a Transformer for Training — an explorable explanationezyang.github.io
- CS231n Deep Learning for Computer Visioncs231n.github.io
- [1709.05011] ImageNet Training in Minutesarxiv-vanity.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- 5D parallelism in LLM training - gdymind's Bloggdymind.com