Scaling Laws Do Not Scale
Recent work has advocated for training AI models on ever-larger datasets, arguing that as the size of a dataset increases, the performance of a model trained on that dataset will correspondingly increase (referred to as “scaling laws”). In this paper, we draw on literature from the social sciences and machine learning to critically interrogate these claims. We argue that this scaling law relationship depends on metrics used to measure performance that may not correspond with how different groups of people perceive the quality of models’ output. As the size of datasets used to train large AI models grows and AI systems impact ever larger groups of people, the number of distinct communities represented in training or evaluation datasets grows. It is thus even more likely that communities represented in datasets may have values or preferences not reflected in (or at odds with) the metrics used to evaluate model performance in scaling laws. Different communities may also have values in ten
Abstract Recent work has advocated for training AI models on ever-larger datasets, arguing that as the size of a dataset increases, the performance of a model trained on that dataset will correspondingly increase (referred to as “scaling laws”). In this paper, we draw on literature from the social sciences and machine learning to critically interrogate these claims. We argue that this scaling law relationship depends on metrics used to measure performance that may not correspond with how different groups of people perceive the quality of models’ output. As the size of datasets used to train…
saved by
related reading
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- The Scaling Hypothesis · Gwern.netgwern.net
- Scaling is subtler than it seemsberen.io
- On neural scaling and the quanta hypothesisericjmichaud.com
- Scaling: The State of Play in AIoneusefulthing.org
- Will scaling work?dwarkeshpatel.com
- AI scaling mythsnormaltech.ai
- A statistical approach to model evaluations \ Anthropicanthropic.com
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- [2405.10938] Observational Scaling Laws and the Predictability of Language Model Performancearxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Ilya Sutskever — We're moving from the age of scaling to the age of researchdwarkesh.com