flâneur

LLM Engineer's Almanac - Advisor | Modal

modal.com · 1,166 words · saved by 1 readers

A simple tool for estimating the throughput and latency of LLM engines

I want to serve with I expect on average Clients should receive in under 95% of the time I want to see configuration Metric: Aggregate: Frequently Asked Questions What is this? How do I use it? This interactive chart indicates the per-replica throughput and client-side latency you can expect when running open weights language models on open source inference engines, in particular on Modal. Select a workload (model, tokens in and out), set a latency objective, and indicate whether you want to see all configurations or just the one that got the best throughput or best latency.…

saved by

related reading