flâneur — a map of the web's best reading

Strangely, Matrix Multiplications on GPUs Run Faster When Given "Predictable" Data! [short]

thonking.ai · 1,884 words · saved by 3 readers

Great minds discuss flops per watt.

Strangely, Matrix Multiplications on GPUs Run Faster When Given "Predictable" Data! [short] Great minds discuss flops per watt. Horace He Apr 29, 2024 170 21 10 Share It’s 2022. I check out this cool new project, CUTLASS , with very fast matmuls. I take a large matmul, 8192 x 8192 x 8192, and benchmark it in PyTorch, which calls CuBLAS. python mm_bench.py > CuBLAS: 258 Teraflops Not bad, 83% flop utilization. Now let’s check out Cutlass’s performance using their profiler. ./cutlass_profiler --operation=Gemm --m=8192 --n=8192 --k=8192 > CUTLASS: 288 Teraflops !!! 10% higher perf? That’s incredi

Explore this link on the map →

saved by

related reading