flâneur — a map of the web's best reading

Massively Scale Your Deep Learning Training with NCCL 2.4 | NVIDIA Technical Blog

developer.nvidia.com · 1,617 words · saved by 1 readers

Imagine using tens of thousands of GPUs to train your neural network. Using multiple GPUs to train neural networks has become quite common with all deep learning frameworks, providing optimized…

Massively Scale Your Deep Learning Training with NCCL 2.4 | NVIDIA Technical Blog Technical Blog Subscribe Related Resources Data Science Massively Scale Your Deep Learning Training with NCCL 2.4 Feb 04, 2019 By Sylvain Jeaugey Like Discuss (1) L T F R E AI-Generated Summary Like Dislike NCCL 2.4 introduces double binary trees, which offer full bandwidth and logarithmic latency for allreduce operations, enabling good performance on small and medium size operations. The double binary tree algorithm significantly improves latency, with up to 180x improvement at 24,576 GPUs on the Summit supercom

Explore this link on the map →

related reading