[2507.10618] Compute Requirements for Algorithmic Innovation in Frontier AI Models
Abstract:Algorithmic innovation in the pretraining of large language models has driven a massive reduction in the total compute required to reach a given level of capability. In this paper we empirically investigate the compute requirements for developing algorithmic innovations. We catalog 36 pre-training algorithmic innovations used in Llama 3 and DeepSeek-V3. For each innovation we estimate both the total FLOP used in development and the FLOP/s of the hardware utilized. Innovations using significant resources double in their requirements each year. We then use this dataset to investigate the effect of compute caps on innovation. Our analysis suggests that compute caps alone are unlikely to dramatically slow AI algorithmic progress. Even stringent compute caps -- such as capping total operations to the compute used to train GPT-2 or capping hardware capacity to 8 H100 GPUs -- could still have allowed for half of the cataloged innovations.
[2507.10618] Compute Requirements for Algorithmic Innovation in Frontier AI Models Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Machine Learning arXiv:2507.10618 (cs) [Submitted on 13 Jul 2025] Title: Compute Requirements for Algorithmic Innovation in Frontier AI Models Authors: Peter Barnett View a PDF of the paper titled Compute Requirements for Algorithmic Innovation in Frontier AI Models, by Peter Barnett View PDF HTML (experimental) Abstract: Algorithmic innovation in the p
Explore this link on the map →related reading
- My picture of the present in AI — LessWronglesswrong.com
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- AI progress is about to speed up | Epoch AIepoch.ai
- Algorithmic Improvement Is Probably Faster Than Scaling Now — LessWronglesswrong.com
- Composer2.pdfcursor.com
- AI in 2025: gestalt — LessWronglesswrong.com
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- Navigating the High Cost of AI Compute | Andreessen Horowitza16z.com
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / Xx.com
- The least understood driver of AI progress | Epoch AIepoch.ai
- Fermi estimate of future training runsdanieldewey.net
- [2403.05812] Algorithmic progress in language modelsarxiv.org