LLM Engineer's Almanac - Advisor | Modal
modal.com · 1,166 words · saved by 1 readers
A simple tool for estimating the throughput and latency of LLM engines
I want to serve with I expect on average Clients should receive in under 95% of the time I want to see configuration Metric: Aggregate: Frequently Asked Questions What is this? How do I use it? This interactive chart indicates the per-replica throughput and client-side latency you can expect when running open weights language models on open source inference engines, in particular on Modal. Select a workload (model, tokens in and out), set a latency objective, and indicate whether you want to see all configurations or just the one that got the best throughput or best latency.…
saved by
related reading
- Wafer | Continual Inferencewafer.ai
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- LLM Engineer's Almanac - Workloads | Modalmodal.com
- LLM Inference Handbookhandbook.modular.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Localmaxxing - Local LLM Inference Speed Testslocalmaxxing.com
- LLM Engineer's Almanac - Executive Summary | Modalmodal.com
- vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention | vLLM Blogblog.vllm.ai
- Inside vLLM: Anatomy of a High-Throughput LLM Inference Systemvllm.ai
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Productizing Large Language Modelsblog.replit.com