Optimally allocating compute between inference and training | Epoch AI
epoch.ai · 2,104 words · saved by 1 readers
AI labs should spend comparable resources on training and inference, assuming they can flexibly balance compute between the two to maintain performance.
Introduction Sam Altman recently claimed that OpenAI currently generates around 100 billion tokens per day, or about 36 trillion tokens per year. Given that modern language models are trained on the order of 10 trillion tokens and tokens seen during training are around three times more expensive1 compared to tokens seen or generated during inference, a naive analysis2 suggests OpenAI’s annual inference costs are on the same order as their annual model training costs. This seems like an odd coincidence at first: why should one of these not completely dominate the other? However, there’s a…
saved by
related reading
- Optimally allocating compute between inference and training | Epoch AIepochai.org
- Trading off compute in training and inference | Epoch AIepochai.org
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Spending Inference Time - Kevin Lukevinlu.ai
- OpenAI’s Strawberry and inference scaling lawsinterconnects.ai
- Navigating the High Cost of AI Compute | Andreessen Horowitza16z.com
- Compute Allocation for AI Discovery and Searchrybindmitry.github.io
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- My picture of the present in AI — LessWronglesswrong.com
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com