One weird trick for parallelizing convolutional neural networks | PDF
arxiv.org · 3,960 words · saved by 1 readers
N/A
One weird trick for parallelizing convolutional neural networks Alex Krizhevsky Google Inc. akrizhevsky@google.com April 29, 2014 arXiv:1404.5997v2 [cs.NE] 26 Apr 2014…
related reading
- How to Parallelize Deep Learning on GPUs Part 1/2: Data Parallelism - Tim Dettmerstimdettmers.com
- Pipeline-Parallelism: Distributed Training via Model Partitioningsiboehm.com
- CS231n Deep Learning for Computer Visioncs231n.github.io
- Paradigms of Parallelism | Colossal-AIcolossalai.org
- Parallelism in Distributed Deep Learning · Better Tomorrow with Computer Scienceinsujang.github.io
- CS231n Deep Learning for Computer Visioncs231n.github.io
- A Recipe for Training Neural Networkskarpathy.github.io
- How to Parallelize a Transformer for Training — an explorable explanationezyang.github.io
- [1706.02677] Accurate, Large Minibatch SGD: Training ImageNet in 1 Hourarxiv-vanity.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- The Little Book of Deep Learningfleuret.org
- Convolutional Neural Networks, Explained | Towards Data Sciencetowardsdatascience.com