✳flâneur — a map of the web's best reading
LLM Inference Performance Engineering: Best Practices | Databricks Blog
databricks.com · 3,725 words · saved by 5 readers
In this blog post, t
LLM Inference Performance Engineering: Best Practices | Databricks Blog Skip to main content In this blog post, the MosaicML engineering team shares best practices for how to capitalize on popular open source large language models (LLMs) for production usage. We also provide guidelines for deploying inference services built around these models to help users in their selection of models and deployment hardware. We have worked with multiple PyTorch-based backends in production; these guidelines are drawn from our experience with FasterTransformers, vLLM, NVIDIA's soon-to-be-released TensorRT-LLM
Explore this link on the map →saved by
related reading
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Optimizing inference · Hugging Facehuggingface.co
- LLM Inference Economics from First Principlestensoreconomics.com
- How LLM Inference Worksarpitbhayani.me
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- A guide to LLM inference and performancebaseten.co
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- How is LLaMa.cpp possible?finbarr.ca
- Inference characteristics of Llama · Cursorcursor.com