My 2 cents on Fusing GEMM + Top-K + Softmax on SM100 | Subho's research at your service 🫡
Mixture-of-Experts (MoE) routing is one of the most latency-sensitive operations in modern LLMs. Every forward pass computes a routing score matrix, selects ...
My 2 cents on Fusing GEMM + Top-K + Softmax on SM100 – Subho's research at your service 🫡 My 2 cents on Fusing GEMM + Top-K + Softmax on SM100 23 Mar, 2026 Mixture-of-Experts (MoE) routing is one of the most latency-sensitive operations in modern LLMs. Every forward pass computes a routing score matrix, selects the top-K experts, and softmax-normalises the weights before dispatching tokens. At inference scale this happens millions of times per second. Shaving microseconds here matters. This post walks through two implementations of a fused GEMM + Top-K + Softmax kernel targeting NVIDIA's Blac
related reading
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Mixture-of-Kittens: our open-source MoE megakernel for NVL72scursor.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- We reverse-engineered Flash Attention 4modal.com
- Grouped GEMM for Imbalanced Experts on Blackwell: A WIP Workloggauravjain.bearblog.dev
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- Crafting Efficient Kernels with Epilogue Fusionblog.fal.ai
- CUTLASS: Fast Linear Algebra in CUDA C++ | NVIDIA Technical Blogdeveloper.nvidia.com
- Worklog: Optimising GEMM on NVIDIA H100 for cuBLAS-like Performance (WIP) – Hamza's Bloghamzaelshafie.bearblog.dev
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev