Will we run out of ML data? Evidence from projecting dataset size trends
Based on our previous analysis of trends in dataset size, we project the growth of dataset size in the language and vision domains. We explore the limits of this trend by estimating the total stock of available unlabeled data over the next decades.
Will we run out of ML data? Projecting dataset size trends | Epoch AI Our projections predict that we will have exhausted the stock of low-quality language data by 2030 to 2050, high-quality language data before 2026, and vision data by 2030 to 2060. This might slow down ML progress. All of our conclusions rely on the unrealistic assumptions that current trends in ML data usage and production will continue and that there will be no major innovations in data efficiency. Relaxing these and other assumptions would be promising future work. Figure 1: ML data consumption and data production trends
related reading
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- Will scaling work?dwarkeshpatel.com
- Researchers warn we could run out of data to train AI by 2026. What then?theconversation.com
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- A Stargate for Data - by willdepue - Will DePuewilldepue.substack.com
- chinchilla's wild implications — AI Alignment Forumalignmentforum.org
- will depue on X: "A Stargate for Data Labs are on a trajectory towards >$100B/year of data spend by 2030. As we begin the trillion-dollar compute project, we need to think about the equivalent civilizational-scale effort for the other core x.com
- chinchilla's wild implications — LessWronglesswrong.com
- AI scaling mythsnormaltech.ai
- Most Algorithmic Progress is Data Progressberen.io
- Can AI scaling continue through 2030? | Epoch AIepoch.ai
- Fermi estimate of future training runsdanieldewey.net