My 2 cents on Fusing GEMM + Top-K + Softmax on SM100 | Subho's research at your service 🫡
Mixture-of-Experts (MoE) routing is one of the most latency-sensitive operations in modern LLMs. Every forward pass computes a routing score matrix, selects ...
My 2 cents on Fusing GEMM + Top-K + Softmax on SM100 – Subho's research at your service 🫡 My 2 cents on Fusing GEMM + Top-K + Softmax on SM100 23 Mar, 2026 Mixture-of-Experts (MoE) routing is one of the most latency-sensitive operations in modern LLMs. Every forward pass computes a routing score matrix, selects the top-K experts, and softmax-normalises the weights before dispatching tokens. At inference scale this happens millions of times per second. Shaving microseconds here matters. This post walks through two implementations of a fused GEMM + Top-K + Softmax kernel targeting NVIDIA's Blac
Explore this link on the map →related reading
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- Crafting Efficient Kernels with Epilogue Fusionblog.fal.ai
- Worklog: Optimising GEMM on NVIDIA H100 for cuBLAS-like Performance (WIP) – Hamza's Bloghamzaelshafie.bearblog.dev
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- CUTLASS: Fast Linear Algebra in CUDA C++ | NVIDIA Technical Blogdeveloper.nvidia.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- 1.5x faster MoE training with custom MXFP8 kernels · Cursorcursor.com