flâneur — a map of the web's best reading

My 2 cents on Fusing GEMM + Top-K + Softmax on SM100 | Subho's research at your service 🫡

ighoshsubho.bearblog.dev · 3,388 words · saved by 1 readers

Mixture-of-Experts (MoE) routing is one of the most latency-sensitive operations in modern LLMs. Every forward pass computes a routing score matrix, selects ...

My 2 cents on Fusing GEMM + Top-K + Softmax on SM100 – Subho's research at your service 🫡 My 2 cents on Fusing GEMM + Top-K + Softmax on SM100 23 Mar, 2026 Mixture-of-Experts (MoE) routing is one of the most latency-sensitive operations in modern LLMs. Every forward pass computes a routing score matrix, selects the top-K experts, and softmax-normalises the weights before dispatching tokens. At inference scale this happens millions of times per second. Shaving microseconds here matters. This post walks through two implementations of a fused GEMM + Top-K + Softmax kernel targeting NVIDIA's Blac

Explore this link on the map →

related reading