Model Runner V2: A Modular and Faster Core for vLLM | vLLM Blog
We are excited to announce Model Runner V2 (MRV2), a ground-up re-implementation of the vLLM model runner. MRV2 delivers a cleaner, more modular, and more effic
Table of Contents We are excited to announce Model Runner V2 (MRV2) , a ground-up re-implementation of the vLLM model runner. MRV2 delivers a cleaner, more modular, and more efficient execution core—with no API changes . The goal is simple: better code and better performance. Like the vLLM V1 release last year, this is an architectural upgrade driven by hard-earned lessons from vLLM's large user base and feedback from the community. We revisited persistent batching, async scheduling, input preparation, and sampling, then rebuilt the model runner around three core principles: Be modular. Isolat
saved by
related reading
- Composer2.pdfcursor.com
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- Inside vLLM: Anatomy of a High-Throughput LLM Inference Systemvllm.ai
- Introduction to vLLM and PagedAttentionblog.runpod.io
- vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention | vLLM Blogblog.vllm.ai
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- M*: A Modular, Extensible, Serving System for Multimodal Modelsai.stanford.edu
- Speculative Decoding - Deep Dive — ROCm Blogsrocm.blogs.amd.com
- Decoding Speculative Decoding from First Principlesjwlabs.vercel.app
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- Speculative Decoding - philkravphilkrav.com
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com