Scaling is subtler than it seems
Epistemic status: A lot of this is vibes and generally trying to express tacit knowledge from ML experience and training models. However there are a huge number of details here and the field of ML changes incredibly rapidly. This is an oversimplified, but hopefully still useful picture, and may well...
Epistemic status: A lot of this is vibes and generally trying to express tacit knowledge from ML experience and training models. However there are a huge number of details here and the field of ML changes incredibly rapidly. This is an oversimplified, but hopefully still useful picture, and may well be wrong in some details since in a good number of cases we understand things at a practitioner’s level but not super well theoretically. People nowadays often think of scaling laws as obvious and almost trivial – if you have a bigger model and use more compute to train it, then it will get…
saved by
related reading
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- On neural scaling and the quanta hypothesisericjmichaud.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- The Scaling Hypothesis · Gwern.netgwern.net
- Pretraining progress is mostly coming from datadwarkesh.com
- Ilya Sutskever — We're moving from the age of scaling to the age of researchdwarkesh.com
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- Will scaling work?dwarkeshpatel.com
- Scaling: The State of Play in AIoneusefulthing.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai