NVIDIA Tensor Core Evolution: From Volta To Blackwell
In our AI Scaling Laws article from late last year, we discussed how multiple stacks of AI scaling laws have continued to drive the AI industry forward, enabling greater than Moore's Law growth in model capabilities as well as a commensurately rapid reduction in unit token costs. These scaling laws are driven by training and inference optimizations and innovations, but advancements in compute capabilities transcending Moore's Law have also played a critical role. One this front, in the AI Scaling Laws article, we revisited the decades-long debate around compute scaling, recounting the end of Dennard Scaling in the late 2000s as well as the end of classic Moore's Law pace cost per transistor declines by the late 2010s. Despite this, compute capabilities have continued to improve at a rapid pace, with the baton being passed to other technologies such as advanced packaging, 3D stacking, new transistor types and specialized architectures such as the GPU. When it comes to AI and deep learni
NVIDIA Tensor Core Evolution: From Volta To Blackwell Amdahl’s Law, Strong Scaling, Asynchronous Execution, Blackwell, Hopper, Ampere, Turing, Volta, TMA Dylan Patel and Kimbo Chen Jun 23, 2025 ∙ Paid 24 2 Share In our AI Scaling Laws article from late last year , we discussed how multiple stacks of AI scaling laws have continued to drive the AI industry forward, enabling greater than Moore's Law growth in model capabilities as well as a commensurately rapid reduction in unit token costs. These scaling laws are driven by training and inference optimizations and innovations, but advancements in
Explore this link on the map →saved by
related reading
- CVPR2023_eff_tutorial_molchanov.pdfnvlabs.github.io
- Worklog: Optimising GEMM on NVIDIA H100 for cuBLAS-like Performance (WIP) – Hamza's Bloghamzaelshafie.bearblog.dev
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- GPU Performance Background User's Guide - NVIDIA Docsdocs.nvidia.com
- H100 vs GB200 NVL72 Training Benchmarks – Power, TCO, and Reliability Analysis, Software Improvement Over Time – SemiAnalysissemianalysis.com
- How to Think About TPUs | How To Scale Your Modeljax-ml.github.io
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Implementing a fast Tensor Core matmul on the Ada Architecture | spatters.caspatters.ca
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- Touching the Elephant - TPUs | Consider the Bulldogconsiderthebulldog.com