Will we run out of ML data? Evidence from projecting dataset size trends
Based on our previous analysis of trends in dataset size, we project the growth of dataset size in the language and vision domains. We explore the limits of this trend by estimating the total stock of available unlabeled data over the next decades.
Will we run out of ML data? Projecting dataset size trends | Epoch AI Our projections predict that we will have exhausted the stock of low-quality language data by 2030 to 2050, high-quality language data before 2026, and vision data by 2030 to 2060. This might slow down ML progress. All of our conclusions rely on the unrealistic assumptions that current trends in ML data usage and production will continue and that there will be no major innovations in data efficiency. Relaxing these and other assumptions would be promising future work. Figure 1: ML data consumption and data production trends
Explore this link on the map →related reading
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- Researchers warn we could run out of data to train AI by 2026. What then?theconversation.com
- chinchilla's wild implications — AI Alignment Forumalignmentforum.org
- Can AI scaling continue through 2030? | Epoch AIepoch.ai
- chinchilla's wild implications — LessWronglesswrong.com
- A Stargate for Data - by willdepue - Will DePuewilldepue.substack.com
- My picture of the present in AI — LessWronglesswrong.com
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- Fermi estimate of future training runsdanieldewey.net
- What will GPT-2030 look like? — AI Alignment Forumalignmentforum.org
- Dario Amodei — "We are near the end of the exponential"dwarkesh.com
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraintsarxiv.org