LLM inference economics from first principles
The main product LLM companies offer these days is access to their models via an API, and the key question that will determine the profitability they can enjoy is the inference cost structure.
LLM Inference Economics from First Principles Piotr Mazurek and Felix Gabriel May 14, 2025 122 11 17 Share The main product LLM companies offer these days is access to their models via an API, and the key question that will determine the profitability they can enjoy is the inference cost structure. In this text we will explain where the cost of serving/hosting LLMs comes from, how many tokens can be produced by a GPU, and why this is the case. We will build a (simplified) world model of LLM inference arithmetics, based on the popular open-source model-LLama 3.3. The goal is to develop an accur
Explore this link on the map →related reading
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- How LLM Inference Worksarpitbhayani.me
- Inference characteristics of Llama · Cursorcursor.com
- A guide to LLM inference and performancebaseten.co
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- How is LLaMa.cpp possible?finbarr.ca
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- Optimizing inference · Hugging Facehuggingface.co
- Inference characteristics of Llama · Cursorcursor.com