The Scaling Hypothesis · Gwern.net
On GPT-3: meta-learning, scaling, implications, and deep theory. The scaling hypothesis: neural nets absorb data & compute, generalizing and becoming more Bayesian as problems get harder, manifesting new abilities even at trivial-by-global-standards-scale. The deep learning revolution has begun as foretold.
--- title: "The Scaling Hypothesis" description: "On GPT-3: meta-learning, scaling, implications, and deep theory. The scaling hypothesis: neural nets absorb data & compute, generalizing and becoming more Bayesian as problems get harder, manifesting new abilities even at trivial-by-global-standards-scale. The deep learning revolution has begun as foretold." thumbnail: /doc/ai/nn/transformer/gpt/2020-brown-gpt3-figure13-meanperformancescalingcurve.png thumbnail-text: "Figure 1.3 from Brown et al 2020 (OpenAI, GPT-3), showing roughly log-scaling of GPT-3 parameter/compute size vs benchmark perfo
related reading
- The Scaling Hypothesis · Gwern.netgwern.net
- On neural scaling and the quanta hypothesisericjmichaud.com
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- Just Ask for Generalization | Eric Jangevjang.com
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- Will scaling work?dwarkeshpatel.com
- Scaling is subtler than it seemsberen.io
- AI progress is about to speed up | Epoch AIepoch.ai
- What will GPT-2030 look like? — AI Alignment Forumalignmentforum.org
- Extrapolating GPT-N performance — AI Alignment Forumalignmentforum.org