New Scaling Laws for Large Language Models - LessWrong
On March 29th, DeepMind published a paper, "Training Compute-Optimal Large Language Models", that shows that essentially everyone -- OpenAI, DeepMind, Microsoft, etc. -- has been training large langu…
x New Scaling Laws for Large Language Models — LessWrong GPT Language Models (LLMs) Machine Learning (ML) AI Frontpage 246 New Scaling Laws for Large Language Models by 1a3orn 1st Apr 2022 AI Alignment Forum 6 min read 22 246 Ω 71 On March 29th, DeepMind published a paper, "Training Compute-Optimal Large Language Models" , that shows that essentially everyone -- OpenAI, DeepMind, Microsoft, etc. -- has been training large language models with a deeply suboptimal use of compute. Following the new scaling laws that they propose for the optimal use of compute, DeepMind trains a new, 70-billion pa
related reading
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- The Scaling Hypothesis · Gwern.netgwern.net
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- chinchilla's wild implications — AI Alignment Forumalignmentforum.org
- chinchilla's wild implications — LessWronglesswrong.com
- On neural scaling and the quanta hypothesisericjmichaud.com
- Scaling is subtler than it seemsberen.io
- [2001.08361] Scaling Laws for Neural Language Modelsarxiv.org
- Chinchillaarxiv.org
- Fermi estimate of future training runsdanieldewey.net
- 2404.10102v1arxiv.org
- Will scaling work?dwarkeshpatel.com