Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. We introduce Mirage Persistent Kernel (MPK), the first compiler and runtime system that automatically transforms multi-GPU model inference into a single high-performance mega-kernel. MPK introduces an SM-level graph representation that captures data dependencies at the granularity of individual streaming multiprocessors (SMs), enabling cross-operator software pipelining, fine-grained kernel overlap, and other previously infeasible GPU optimizations. The MPK compiler lowers tensor programs into highly optimized SM-level task graphs and generates optimized CUDA implementations for all tasks, while the MPK in-kernel parallel runtime executes these tasks within a single mega-kernel using decentralized scheduling across SMs. Together, these components provide
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs Xinhao Cheng 1,∗ Zhihao Zhang 1,∗ Yu Zhou 1,∗ Jianan Ji 1,∗ Jinchen Jiang 2 Zepeng Zhao 1 Ziruo Xiao 1 Zihao Ye 3 Yingyi Huang 3 Ruihang Lai 1 Hongyi Jin 1 Bohan Hou 1 Mengdi Wu 1 Yixin Dong 1 Anthony Yip 1 Zihao Ye 4 Songting Wang 1 Wenqin Yang 5 Xupeng Miao 6 Tianqi Chen 1,3 Zhihao Jia 1 Carnegie Mellon University 1 Tsinghua University 2 NVIDIA 3 University of Michigan 4 Independent Researcher 5 Purdue University 6 Abstract We introduce Mirage Persistent Kernel (MPK), the first compiler and runtime system t
Explore this link on the map →saved by
related reading
- Compiling Models to Megakernels - by Luminal and Joe Fiotiblog.luminal.com
- Real-time LLM Inference on Standard Datacenter GPUs (3,000 tokens/s per request)blog.kog.ai
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- [2410.20399] ThunderKittens: Simple, Fast, and Adorable AI Kernelsarxiv.org
- TPU Deep Divehenryhmko.github.io
- PiTorch: ML on Baremetal Raspberry Pis | projectsmasonjwang.com
- A friendly introduction to machine learning compilers and optimizershuyenchip.com
- How to Think About GPUs | How To Scale Your Modeljax-ml.github.io
- Pipeline-Parallelism: Distributed Training via Model Partitioningsiboehm.com