Tips for Optimizing GPU Performance Using Tensor Cores | NVIDIA Technical Blog
Our most popular question is "What can I do to get great GPU performance for deep learning?" We’ve recently published a detailed Deep Learning Performance Guide to help answer this question.
Tips for Optimizing GPU Performance Using Tensor Cores | NVIDIA Technical Blog Technical Blog Subscribe Related Resources Data Science Tips for Optimizing GPU Performance Using Tensor Cores Jun 10, 2019 By Valerie Sarge , Michael Andersch , Lynsey Fabel , Paulius Micikevicius and John Tran Like Discuss (15) L T F R E AI-Generated Summary Like Dislike To activate Tensor Cores on NVIDIA GPUs, parameters such as batch size and number of inputs and outputs should be divisible by 8 for FP16 data or 16 for INT8 data, as seen in the Transformer neural network's projection layer where padding voc
Explore this link on the map →related reading
- Linear/Fully-Connected Layers User's Guide - NVIDIA Docsdocs.nvidia.com
- CVPR2023_eff_tutorial_molchanov.pdfnvlabs.github.io
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- Get Started With Deep Learning Performance - NVIDIA Docsdocs.nvidia.com
- GPU Performance Background User's Guide - NVIDIA Docsdocs.nvidia.com
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- NVIDIA Tensor Core Evolution: From Volta To Blackwellnewsletter.semianalysis.com
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- How To Scale Your Modeljax-ml.github.io
- Making Deep Learning go Brrrr From First Principleshorace.io
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Outperforming cuBLAS on H100: a Worklogcudaforfun.substack.com