chinchilla's wild implications — AI Alignment Forum
(Colab notebook here.) • This post is about language model scaling laws, specifically the laws derived in the DeepMind paper that introduced Chinchilla.[1] …
x chinchilla's wild implications — AI Alignment Forum Best of LessWrong 2022 Scaling Laws Language Models (LLMs) Machine Learning (ML) DeepMind AI Frontpage 97 chinchilla's wild implications by nostalgebraist 31st Jul 2022 13 min read 129 97 ( Colab notebook here.) This post is about language model scaling laws, specifically the laws derived in the DeepMind paper that introduced Chinchilla. [1] The paper came out a few months ago, and has been discussed a lot, but some of its implications deserve more explicit notice in my opinion. In particular: Data, not size, is the currently active constra
related reading
- chinchilla's wild implications — LessWronglesswrong.com
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- Will scaling work?dwarkeshpatel.com
- Chinchillaarxiv.org
- The Scaling Hypothesis · Gwern.netgwern.net
- Scaling is subtler than it seemsberen.io
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraintsarxiv.org
- Scaling: The State of Play in AIoneusefulthing.org
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- [2509.14786] Pre-training under infinite computearxiv.org