chinchilla's wild implications — AI Alignment Forum
(Colab notebook here.) • This post is about language model scaling laws, specifically the laws derived in the DeepMind paper that introduced Chinchilla.[1] …
x chinchilla's wild implications — AI Alignment Forum Best of LessWrong 2022 Scaling Laws Language Models (LLMs) Machine Learning (ML) DeepMind AI Frontpage 97 chinchilla's wild implications by nostalgebraist 31st Jul 2022 13 min read 129 97 ( Colab notebook here.) This post is about language model scaling laws, specifically the laws derived in the DeepMind paper that introduced Chinchilla. [1] The paper came out a few months ago, and has been discussed a lot, but some of its implications deserve more explicit notice in my opinion. In particular: Data, not size, is the currently active constra
Explore this link on the map →related reading
- chinchilla's wild implications — LessWronglesswrong.com
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- The Scaling Hypothesis · Gwern.netgwern.net
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- On neural scaling and the quanta hypothesisericjmichaud.com
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraintsarxiv.org
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com
- [2203.15556] Training Compute-Optimal Large Language Modelsarxiv.org
- Thoughts on the Alignment Implications of Scaling Language Models | Leo Gaobmk.sh