The Scaling Hypothesis · Gwern.net
On GPT-3: meta-learning, scaling, implications, and deep theory. The scaling hypothesis: neural nets absorb data & compute, generalizing and becoming more Bayesian as problems get harder, manifesting new abilities even at trivial-by-global-standards-scale. The deep learning revolution has begun as foretold.
--- title: "The Scaling Hypothesis" description: "On GPT-3: meta-learning, scaling, implications, and deep theory. The scaling hypothesis: neural nets absorb data & compute, generalizing and becoming more Bayesian as problems get harder, manifesting new abilities even at trivial-by-global-standards-scale. The deep learning revolution has begun as foretold." thumbnail: /doc/ai/nn/transformer/gpt/2020-brown-gpt3-figure13-meanperformancescalingcurve.png thumbnail-text: "Figure 1.3 from Brown et al 2020 (OpenAI, GPT-3), showing roughly log-scaling of GPT-3 parameter/compute size vs benchmark perfo
saved by
- Elizabeth Qiu
- Rajan Agarwal
- Ratan Kaliani
- Aryan Naik
- Katherine Slattery
- Vedant Nair
- KL Yap
- Lydia Nottingham
- Yash Dani
- Yudhister Joel Kumar
- Timothy Kostolansky
- Eric Huang
related reading
- The Scaling Hypothesis · Gwern.netgwern.net
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- On neural scaling and the quanta hypothesisericjmichaud.com
- Ilya Sutskever — We're moving from the age of scaling to the age of researchdwarkesh.com
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- What will GPT-2030 look like? — AI Alignment Forumalignmentforum.org
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- Will scaling work?dwarkeshpatel.com
- Fermi estimate of future training runsdanieldewey.net
- Just Ask for Generalization | Eric Jangevjang.com