Model Runner V2: A Modular and Faster Core for vLLM | vLLM Blog
We are excited to announce Model Runner V2 (MRV2), a ground-up re-implementation of the vLLM model runner. MRV2 delivers a cleaner, more modular, and more effic
Table of Contents We are excited to announce Model Runner V2 (MRV2) , a ground-up re-implementation of the vLLM model runner. MRV2 delivers a cleaner, more modular, and more efficient execution core—with no API changes . The goal is simple: better code and better performance. Like the vLLM V1 release last year, this is an architectural upgrade driven by hard-earned lessons from vLLM's large user base and feedback from the community. We revisited persistent batching, async scheduling, input preparation, and sampling, then rebuilt the model runner around three core principles: Be modular. Isolat
Explore this link on the map →saved by
related reading
- Composer2.pdfcursor.com
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention | vLLM Blogblog.vllm.ai
- Speculative Decoding - Deep Dive — ROCm Blogsrocm.blogs.amd.com
- 2025 LLM Year in Review – karpathykarpathy.bearblog.dev
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- Speculative Decoding - philkravphilkrav.com
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- How to make LLMs go fastvgel.me
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- API Reference — TensorRT LLMnvidia.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com