✳flâneur — a map of the web's best reading
Strangely, Matrix Multiplications on GPUs Run Faster When Given "Predictable" Data! [short]
thonking.ai · 1,884 words · saved by 3 readers
Great minds discuss flops per watt.
Strangely, Matrix Multiplications on GPUs Run Faster When Given "Predictable" Data! [short] Great minds discuss flops per watt. Horace He Apr 29, 2024 170 21 10 Share It’s 2022. I check out this cool new project, CUTLASS , with very fast matmuls. I take a large matmul, 8192 x 8192 x 8192, and benchmark it in PyTorch, which calls CuBLAS. python mm_bench.py > CuBLAS: 258 Teraflops Not bad, 83% flop utilization. Now let’s check out Cutlass’s performance using their profiler. ./cutlass_profiler --operation=Gemm --m=8192 --n=8192 --k=8192 > CUTLASS: 288 Teraflops !!! 10% higher perf? That’s incredi
Explore this link on the map →saved by
related reading
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- Making Deep Learning go Brrrr From First Principleshorace.io
- “This Kernel Was Faster Yesterday” — In Pursuit of High-Fidelity GPU Kernel Benchmarkingstandardkernel.com
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- GPU Performance Background User's Guide - NVIDIA Docsdocs.nvidia.com
- Mini Project: GPU Accelerated Matrix Multiplication (almost) like cuBLAS0mean1sigma.com
- All About Rooflines | How To Scale Your Modeljax-ml.github.io
- Outperforming cuBLAS on H100: a Worklogcudaforfun.substack.com
- Outperforming cuBLAS on H100: a Worklogcudaforfun.substack.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly