New Scaling Laws for Large Language Models - LessWrong
On March 29th, DeepMind published a paper, "Training Compute-Optimal Large Language Models", that shows that essentially everyone -- OpenAI, DeepMind, Microsoft, etc. -- has been training large langu…
x New Scaling Laws for Large Language Models — LessWrong GPT Language Models (LLMs) Machine Learning (ML) AI Frontpage 246 New Scaling Laws for Large Language Models by 1a3orn 1st Apr 2022 AI Alignment Forum 6 min read 22 246 Ω 71 On March 29th, DeepMind published a paper, "Training Compute-Optimal Large Language Models" , that shows that essentially everyone -- OpenAI, DeepMind, Microsoft, etc. -- has been training large language models with a deeply suboptimal use of compute. Following the new scaling laws that they propose for the optimal use of compute, DeepMind trains a new, 70-billion pa
Explore this link on the map →related reading
- The Scaling Hypothesis · Gwern.netgwern.net
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- chinchilla's wild implications — AI Alignment Forumalignmentforum.org
- chinchilla's wild implications — LessWronglesswrong.com
- [2001.08361] Scaling Laws for Neural Language Modelsarxiv.org
- On neural scaling and the quanta hypothesisericjmichaud.com
- How To Scale Your Modeljax-ml.github.io
- Fermi estimate of future training runsdanieldewey.net
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- The Little Book of Deep Learningfleuret.org
- Bits, FLOPS, and Watts: A Systems-Level Perspective of Scaling LLMs — Part 1 | by Asheesh Goja | Mediummedium.com