LLM Inference Performance Engineering: Best Practices | Databricks Blog
databricks.com · 3,725 words · saved by 6 readers
In this blog post, t
LLM Inference Performance Engineering: Best Practices | Databricks Blog Skip to main content In this blog post, the MosaicML engineering team shares best practices for how to capitalize on popular open source large language models (LLMs) for production usage. We also provide guidelines for deploying inference services built around these models to help users in their selection of models and deployment hardware. We have worked with multiple PyTorch-based backends in production; these guidelines are drawn from our experience with FasterTransformers, vLLM, NVIDIA's soon-to-be-released TensorRT-LLM
saved by
related reading
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- Efficient LLM inferencefinbarrtimbers.substack.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- Optimizing inference · Hugging Facehuggingface.co
- LLM Inference Economics from First Principlestensoreconomics.com
- How LLM Inference Worksarpitbhayani.me
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- A guide to LLM inference and performancebaseten.co
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- Inference characteristics of Llama · Cursorcursor.com