[2509.14786] Pre-training under infinite compute
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2509.14786] Pre-training under infinite compute --> Computer Science > Machine Learning arXiv:2509.14786 (cs) [Submitted on 18 Sep 2025] Title: Pre-training under infinite compute Authors: Konwoo Kim , Suhas Kotha , Percy Liang , Tatsunori Hashimoto View a PDF of the paper titled Pre-training under infinite compute, by Konwoo Kim and Suhas Kotha and Percy Liang and Tatsunori Hashimoto View PDF HTML (experimental) Abstract: Since compute grows much faster than web text available for language model pre-training, we ask how one should approach pre-training under fixed data and no compute constra
saved by
related reading
- [2509.14786] Pre-training under infinite computearxiv.org
- Pre-training under infinite computearxiv.org
- >10x More Efficient Pretraining — Magicmagic.dev
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- 10x Data Efficiency - NanoGPT Slowrunqlabs.sh
- Pretraining progress is mostly coming from datadwarkesh.com
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- The Scaling Hypothesis · Gwern.netgwern.net
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Composer2.pdfcursor.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraintsarxiv.org