flâneur

LLM Inference Handbook

handbook.modular.com · 664 words · saved by 1 readers

A practical handbook for engineers building, optimizing, scaling and operating LLM inference systems in production.

LLM Inference Handbook is your technical glossary, guidebook, and reference - all in one. It covers everything you need to know about LLM inference, from core concepts and performance metrics (e.g., Time to First Token and Tokens per Second), to optimization techniques (e.g., continuous batching and prefix caching), GPU architecture, and deployment patterns like BYOC and on-prem. Practical guidance for deploying, scaling, and operating LLMs in production. Explore concepts with interactive calculators, simulators, and visual tools. Boost performance with optimization techniques tailored to…

saved by

related reading