LLM Inference Handbook
handbook.modular.com · 664 words · saved by 1 readers
A practical handbook for engineers building, optimizing, scaling and operating LLM inference systems in production.
LLM Inference Handbook is your technical glossary, guidebook, and reference - all in one. It covers everything you need to know about LLM inference, from core concepts and performance metrics (e.g., Time to First Token and Tokens per Second), to optimization techniques (e.g., continuous batching and prefix caching), GPU architecture, and deployment patterns like BYOC and on-prem. Practical guidance for deploying, scaling, and operating LLMs in production. Explore concepts with interactive calculators, simulators, and visual tools. Boost performance with optimization techniques tailored to…
saved by
related reading
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- GenAI Handbookgenai-handbook.github.io
- Optimizing inference · Hugging Facehuggingface.co
- LLM Engineer's Almanac - Advisormodal.com
- Together AI | The AI Native Cloudtogether.ai
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- A Guide to AI Inference Engineering - ByteByteGo Newsletterblog.bytebytego.com
- How LLM Inference Worksarpitbhayani.me
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com