Crafting Efficient Kernels with Epilogue Fusion
In many ML workloads, a GEMM is followed by small operations like bias, activation, scaling, or type conversion. These ops are cheap in math, but they often cost extra global memory traffic (store GEMM result, read it back, write again). Epilogue fusion is a way to avoid this, we can
In many ML workloads, a GEMM is followed by small operations like bias, activation, scaling, or type conversion. These ops are cheap in math, but they often cost extra global memory traffic (store GEMM result, read it back, write again). Epilogue fusion is a way to avoid this, we can apply these extra ops while the GEMM result is still in registers, right before the final store to global memory. On Hopper and Blackwell, there is also more room to overlap Tensor Core work with other instructions, so doing some extra compute in the epilogue can be even more attractive. Epilogue fusion eliminates
Explore this link on the map →saved by
related reading
- Implementing a fast Tensor Core matmul on the Ada Architecture | spatters.caspatters.ca
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- CUTLASS: Fast Linear Algebra in CUDA C++ | NVIDIA Technical Blogdeveloper.nvidia.com
- My 2 cents on Fusing GEMM + Top-K + Softmax on SM100 – Subho's research at your service 🫡ighoshsubho.bearblog.dev
- Worklog: Optimising GEMM on NVIDIA H100 for cuBLAS-like Performance (WIP) – Hamza's Bloghamzaelshafie.bearblog.dev
- Tiny TPUtinytpu.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Outperforming cuBLAS on H100: a Worklogcudaforfun.substack.com
- Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programsarxiv.org
- Outperforming cuBLAS on H100: a Worklogcudaforfun.substack.com
- Mini Project: GPU Accelerated Matrix Multiplication (almost) like cuBLAS0mean1sigma.com