A history of NVidia Stream Multiprocessor
I spent last week-end getting accustomed to CUDA and SIMT programming. It was a prolific time ending up with a Business Card Raytracer running close to 700x faster[1], from 101s to 150ms. This pleasant experience was a good pretext to spend more time on the topic and learn about the evolution of Nvidia architecture. Thanks to the abundant documentation published over the years by the green team, I was able to go back in time and fast forward though the fascinating evolution of their stream multiprocessors. Visited in this article: Up to 2006, NVidia's GPU design was correlated to the logical stages in the rendering API[2]. The GeForce 7900 GTX, powered by a G71 die is made of three sections dedicated to vertex processing (8 units), fragment generation (24 units), and fragment merging (16 units). The G71. Notice the Z-Cull optimization discarding fragment that would fail the Z test. This correlation forced designers to guess the location of bottlenecks in order to properly balance e
A history of NVidia Stream Multiprocessor FABIEN SANGLARD'S WEBSITE CONTACT RSS DONATE May 2, 2020 A history of NVidia Stream Multiprocessor I spent last week-end getting accustomed to CUDA and SIMT programming. It was a prolific time ending up with a Business Card Raytracer running close to 700x faster [1] , from 101s to 150ms. This pleasant experience was a good pretext to spend more time on the topic and learn about the evolution of Nvidia architecture. Thanks to the abundant documentation published over the years by the green team, I was able to go back in time and fast forward though the
Explore this link on the map →saved by
related reading
- Demystifying GPU Compute Architectures - by Babbagethechipletter.substack.com
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- Execution Model - SLING user documentationdoc.sling.si
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- What is a Streaming Multiprocessor? | GPU Glossarymodal.com
- CUDA C++ Programming Guide (Legacy) — CUDA C++ Programming Guidedocs.nvidia.com
- NVIDIA Tensor Core Evolution: From Volta To Blackwellnewsletter.semianalysis.com
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- Worklog: Optimising GEMM on NVIDIA H100 for cuBLAS-like Performance (WIP) – Hamza's Bloghamzaelshafie.bearblog.dev
- CUDA - Wikipediaen.wikipedia.org
- Outperforming cuBLAS on H100: a Worklogcudaforfun.substack.com