flâneur — a map of the web's best reading

Worklog: Optimising GEMM on NVIDIA H100 for cuBLAS-like Performance (WIP) | Hamza's Blog

hamzaelshafie.bearblog.dev · 22,076 words · saved by 1 readers

Matrix multiplication sits at the core of modern deep learning. Whether it is transformers, CNNs, or even simple MLPs, everything eventually reduces to G...

Worklog: Optimising GEMM on NVIDIA H100 for cuBLAS-like Performance (WIP) – Hamza's Blog Worklog: Optimising GEMM on NVIDIA H100 for cuBLAS-like Performance (WIP) 12 Jan, 2026 🚧 Work in progress. Please reach out on LinkedIn if you spot any mistakes. Introduction Matrix multiplication sits at the core of modern deep learning. Whether it is transformers, CNNs, or even simple MLPs, everything eventually reduces to GEMM. GPUs are built to run this operation at scale, and libraries like cuBLAS set the performance bar with kernels tuned down to the last instruction. In this blog I am rebuilding th

Explore this link on the map →

saved by

related reading