Pre-training under infinite compute
arxiv.org · 8,630 words · saved by 1 readers
N/A
Pre-training under infinite compute Konwoo Kim∞ , Suhas Kotha∞ , Percy Liang, Tatsunori Hashimoto Stanford University Abstract arXiv:2509.14786v1 [cs.LG] 18 Sep 2025 Since compute grows much faster than web text available for language model pre-training, we ask how one…
related reading
- [2509.14786] Pre-training under infinite computearxiv.org
- [2509.14786] Pre-training under infinite computearxiv.org
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- 10x Data Efficiency - NanoGPT Slowrunqlabs.sh
- >10x More Efficient Pretraining — Magicmagic.dev
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- The Scaling Hypothesis · Gwern.netgwern.net
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraintsarxiv.org
- Pretraining progress is mostly coming from datadwarkesh.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Scaling is subtler than it seemsberen.io
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai