10x Data Efficiency - NanoGPT Slowrun - Q
We've achieved 10x data efficiency with NanoGPT Slowrun within a few weeks. An ensemble of 1.8B parameter models (18B total params) trained on 100M tokens matches what would normally require 1B tokens with a standard LM baseline. Data efficiency matters because compute grows much faster than data . Since our current scaling laws require proportional increases in both , intelligence will eventually be bottlenecked by data, not compute. This data efficiency result allows us to improve model performance by scaling with compute rather than with data. A few things worth noting. First, this looks nothing like our current scaling laws. Chinchilla says you should train a ~5M parameter model if you have 100M tokens -- a staggering 3600x difference from what we're doing. Second, 10x data efficiency would've seemed unimaginable to most people, and we got there in ... a few weeks. Here's how. Some of the trends are architectural tweaks without a lot of principles behind them. But a few are princi
We've achieved 10x data efficiency with NanoGPT Slowrun within a few weeks. An ensemble of 1.8B parameter models (18B total params) trained on 100M tokens matches what would normally require 1B tokens with a standard LM baseline. Data efficiency matters because compute grows much faster than data . Since our current scaling laws require proportional increases in both , intelligence will eventually be bottlenecked by data, not compute. This data efficiency result allows us to improve model performance by scaling with compute rather than with data. NanoGPT Slowrun 3.8× data efficiency A few…
saved by
related reading
- How To Scale Your Modeljax-ml.github.io
- [2509.14786] Pre-training under infinite computearxiv.org
- The Scaling Hypothesis · Gwern.netgwern.net
- Pre-training under infinite computearxiv.org
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- [2509.14786] Pre-training under infinite computearxiv.org
- >10x More Efficient Pretraining — Magicmagic.dev
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- Pretraining progress is mostly coming from datadwarkesh.com
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai