flâneur — a map of the web's best reading

Will we run out of ML data? Evidence from projecting dataset size trends

epochai.org · 1,089 words · saved by 1 readers

Based on our previous analysis of trends in dataset size, we project the growth of dataset size in the language and vision domains. We explore the limits of this trend by estimating the total stock of available unlabeled data over the next decades.

Will we run out of ML data? Projecting dataset size trends | Epoch AI Our projections predict that we will have exhausted the stock of low-quality language data by 2030 to 2050, high-quality language data before 2026, and vision data by 2030 to 2060. This might slow down ML progress. All of our conclusions rely on the unrealistic assumptions that current trends in ML data usage and production will continue and that there will be no major innovations in data efficiency. Relaxing these and other assumptions would be promising future work. Figure 1: ML data consumption and data production trends

Explore this link on the map →

related reading