flâneur — a map of the web's best reading

A Bitter Lesson for Data Filtering

arxiv.org · 11,320 words · saved by 1 readers

We investigate data filtering for large model pretraining via new scaling studies that target the high compute, data-scarce regime. In spite of an apparently common belief that filtering data to include only high-quality information is essential, our experiments suggest that with enough compute, the best data filter is no data filter. We find that sufficiently trained large parameter models not only tolerate low-quality and distractor data, but in fact benefit from nominally “poor” data. The standard approach to select pretraining data for language models is to filter text from sources like Common Crawl (CC) (Common Crawl, 2024). It is widely documented that in compute-constrained regimes, where one must train on a subset of CC, different data selection strategies can have a large impact on performance. This is intuitive: all else equal, it seems natural to train on “higher-quality” data. As a result, a large body of research has emerged to tackle the data selection problem, with the g

A Bitter Lesson for Data Filtering Christopher Mohri Department of Computer Science Stanford University xmohri@stanford.edu &John Duchi Departments of Statistics and Electrical Engineering Stanford University jduchi@stanford.edu &Tatsunori Hashimoto Department of Computer Science Stanford University thashim@stanford.edu Abstract We investigate data filtering for large model pretraining via new scaling studies that target the high compute, data-scarce regime. In spite of an apparently common belief that filtering data to include only high-quality information is essential, our experiments sugges

Explore this link on the map →

saved by

related reading