The Scaling Hypothesis · Gwern.net
On GPT-3: meta-learning, scaling, implications, and deep theory. The scaling hypothesis: neural nets absorb data & compute, generalizing and becoming more Bayesian as problems get harder, manifesting new abilities even at trivial-by-global-standards-scale. The deep learning revolution has begun as foretold.
--- title: "The Scaling Hypothesis" description: "On GPT-3: meta-learning, scaling, implications, and deep theory. The scaling hypothesis: neural nets absorb data & compute, generalizing and becoming more Bayesian as problems get harder, manifesting new abilities even at trivial-by-global-standards-scale. The deep learning revolution has begun as foretold." thumbnail: /doc/ai/nn/transformer/gpt/2020-brown-gpt3-figure13-meanperformancescalingcurve.png thumbnail-text: "Figure 1.3 from Brown et al 2020 (OpenAI, GPT-3), showing roughly log-scaling of GPT-3 parameter/compute size vs benchmark perfo
Explore this link on the map →related reading
- The Scaling Hypothesis · Gwern.netgwern.net
- On neural scaling and the quanta hypothesisericjmichaud.com
- Just Ask for Generalization | Eric Jangevjang.com
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- Extrapolating GPT-N performance — AI Alignment Forumalignmentforum.org
- What will GPT-2030 look like? — AI Alignment Forumalignmentforum.org
- AI progress is about to speed up | Epoch AIepoch.ai
- Fermi estimate of future training runsdanieldewey.net
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- Dario Amodei — "We are near the end of the exponential"dwarkesh.com
- Are we in an AI overhang? — LessWronglesswrong.com