[2506.04645] Inference economics of language models
Abstract:We develop a theoretical model that addresses the economic trade-off between cost per token versus serial token generation speed when deploying LLMs for inference at scale. Our model takes into account arithmetic, memory bandwidth, network bandwidth and latency constraints; and optimizes over different parallelism setups and batch sizes to find the ones that optimize serial inference speed at a given cost per token. We use the model to compute Pareto frontiers of serial speed versus cost per token for popular language models.
View PDF HTML (experimental) Abstract:We develop a theoretical model that addresses the economic trade-off between cost per token versus serial token generation speed when deploying LLMs for inference at scale. Our model takes into account arithmetic, memory bandwidth, network bandwidth and latency constraints; and optimizes over different parallelism setups and batch sizes to find the ones that optimize serial inference speed at a given cost per token. We use the model to compute Pareto frontiers of serial speed versus cost per token for popular language models. Subjects: Machine…
saved by
related reading
- Continuous batching from first principleshuggingface.co
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- LLM Inference Economics from First Principlestensoreconomics.com
- How To Scale Your Modeljax-ml.github.io
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Inference characteristics of Llama · Cursorcursor.com
- Small Models Have Arrivedcalv.info
- Optimizing inference · Hugging Facehuggingface.co
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- LLM Engineer's Almanac - Advisormodal.com
- Spending Inference Time - Kevin Lukevinlu.ai