chinchilla's wild implications - LessWrong
The DeepMind paper that introduced Chinchilla revealed that we've been using way too many parameters and not enough data for large language models. T…
x chinchilla's wild implications — LessWrong Best of LessWrong 2022 Scaling Laws Language Models (LLMs) Machine Learning (ML) DeepMind AI Frontpage 425 chinchilla's wild implications by nostalgebraist 31st Jul 2022 AI Alignment Forum 13 min read 129 425 Ω 97 ( Colab notebook here.) This post is about language model scaling laws, specifically the laws derived in the DeepMind paper that introduced Chinchilla. [1] The paper came out a few months ago, and has been discussed a lot, but some of its implications deserve more explicit notice in my opinion. In particular: Data, not size, is the current
Explore this link on the map →related reading
- chinchilla's wild implications — AI Alignment Forumalignmentforum.org
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- The Scaling Hypothesis · Gwern.netgwern.net
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- Norman Mu | The Myth of Data Inefficiency in Large Language Modelsnormanmu.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- [2203.15556] Training Compute-Optimal Large Language Modelsarxiv.org
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraintsarxiv.org
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai