A Bitter Lesson for Data Filtering
We investigate data filtering for large model pretraining via new scaling studies that target the high compute, data-scarce regime. In spite of an apparently common belief that filtering data to include only high-quality information is essential, our experiments suggest that with enough compute, the best data filter is no data filter. We find that sufficiently trained large parameter models not only tolerate low-quality and distractor data, but in fact benefit from nominally “poor” data. The standard approach to select pretraining data for language models is to filter text from sources like Common Crawl (CC) (Common Crawl, 2024). It is widely documented that in compute-constrained regimes, where one must train on a subset of CC, different data selection strategies can have a large impact on performance. This is intuitive: all else equal, it seems natural to train on “higher-quality” data. As a result, a large body of research has emerged to tackle the data selection problem, with the g
A Bitter Lesson for Data Filtering Christopher Mohri Department of Computer Science Stanford University xmohri@stanford.edu &John Duchi Departments of Statistics and Electrical Engineering Stanford University jduchi@stanford.edu &Tatsunori Hashimoto Department of Computer Science Stanford University thashim@stanford.edu Abstract We investigate data filtering for large model pretraining via new scaling studies that target the high compute, data-scarce regime. In spite of an apparently common belief that filtering data to include only high-quality information is essential, our experiments sugges
Explore this link on the map →saved by
related reading
- DataRater: Meta-Learned Dataset Curationarxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraintsarxiv.org
- [2509.14786] Pre-training under infinite computearxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Data Management For Large Language Models: A Surveyarxiv.org
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scalearxiv.org
- Good QC for RL Dataseancai.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io