✳flâneur — a map of the web's best reading
“This Kernel Was Faster Yesterday” — In Pursuit of High-Fidelity GPU Kernel Benchmarking
standardkernel.com · 4,044 words · saved by 1 readers
“This Kernel Was Faster Yesterday” — In Pursuit of High-Fidelity GPU Kernel Benchmarking
“This Kernel Was Faster Yesterday” — In Pursuit of High-Fidelity GPU Kernel Benchmarking ← Blog 24 February 2026 “This Kernel Was Faster Yesterday” — In Pursuit of High-Fidelity GPU Kernel Benchmarking After Salvador Dalí’s The Persistence of Memory TL;DR GPU timing is deceptively hard and highly variable : power limits, thermal state, clock behavior, idle transitions, caching, and measurement methods all matter. High-fidelity evaluation is critical, especially for automated RL systems. In high-value matmul kernels, where even 5% matters, measurement noise can look like real gains and mislead
Explore this link on the map →related reading
- Strangely, Matrix Multiplications on GPUs Run Faster When Given "Predictable" Data! [short]thonking.ai
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- KernelBench v0.1 | Scaling Intelligence Lab at Stanford Universityscalingintelligence.stanford.edu
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- Making Deep Learning go Brrrr From First Principleshorace.io
- GPU Performance Background User's Guide - NVIDIA Docsdocs.nvidia.com
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- CUDA C++ Programming Guide (Legacy) — CUDA C++ Programming Guidedocs.nvidia.com
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- Worklog: Optimising GEMM on NVIDIA H100 for cuBLAS-like Performance (WIP) – Hamza's Bloghamzaelshafie.bearblog.dev