Efficient Memory Management for Large Language Model Serving with PagedAttention
arxiv.org · 5,348 words · saved by 1 readers
N/A
Efficient Memory Management for Large Language Model Serving with PagedAttention Woosuk Kwon1,∗ Zhuohan Li1,∗ Siyuan Zhuang1 Ying Sheng1,2 Lianmin Zheng1 Cody Hao Yu3 Joseph E. Gonzalez1 Hao Zhang4 Ion Stoica1 1 UC Berkeley 2 Stanford University 3 Independent Researcher 4 UC San Diego…
related reading
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- Introduction to vLLM and PagedAttentionblog.runpod.io
- vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention | vLLM Blogblog.vllm.ai
- Inside vLLM: Anatomy of a High-Throughput LLM Inference Systemvllm.ai
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Continuous batching from first principleshuggingface.co
- 2306.14048arxiv.org
- Cache strategies · Hugging Facehuggingface.co
- At the Intersection of LLMs and Kernels - Research Roundupcharlesfrye.github.io
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu