A Bitter Lesson for Data Filtering
We investigate data filtering for large model pretraining via new scaling studies that target the high compute, data-scarce regime. In spite of an apparently common belief that filtering data to include only high-quality information is essential, our experiments suggest that with enough compute, the best data filter is no data filter. We find that sufficiently trained large parameter models not only tolerate low-quality and distractor data, but in fact benefit from nominally “poor” data. The standard approach to select pretraining data for language models is to filter text from sources like Common Crawl (CC) (Common Crawl, 2024). It is widely documented that in compute-constrained regimes, where one must train on a subset of CC, different data selection strategies can have a large impact on performance. This is intuitive: all else equal, it seems natural to train on “higher-quality” data. As a result, a large body of research has emerged to tackle the data selection problem, with the g
A Bitter Lesson for Data Filtering Christopher Mohri Department of Computer Science Stanford University xmohri@stanford.edu &John Duchi Departments of Statistics and Electrical Engineering Stanford University jduchi@stanford.edu &Tatsunori Hashimoto Department of Computer Science Stanford University thashim@stanford.edu Abstract We investigate data filtering for large model pretraining via new scaling studies that target the high compute, data-scarce regime. In spite of an apparently common belief that filtering data to include only high-quality information is essential, our experiments sugges
saved by
related reading
- Shaping capabilities with token-level data filteringarxiv.org
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- [2509.14786] Pre-training under infinite computearxiv.org
- DataRater: Meta-Learned Dataset Curationarxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraintsarxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Scaling is subtler than it seemsberen.io
- Data Management For Large Language Models: A Surveyarxiv.org
- [2509.14786] Pre-training under infinite computearxiv.org
- Pre-training under infinite computearxiv.org
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io