The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale Guilherme Penedo Hynek Kydlíček Loubna Ben allal Anton Lozhkov Margaret Mitchell Colin Raffel Leandro Von Werra Thomas Wolf Hugging Face Abstract The performance of a large language model (LLM) depends heavily on the quality and size of its pretraining dataset. However, the pretraining datasets for state-of-the-art open LLMs like Llama 3 and Mixtral are not publicly available and very little is known about how they were created. In this work, we introduce FineWeb, a 15-trillion token dataset derived from 96 Common Crawl
related reading
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- DataRater: Meta-Learned Dataset Curationarxiv.org
- A Bitter Lesson for Data Filteringarxiv.org
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Shaping capabilities with token-level data filteringarxiv.org
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretrainingdatologyai.com
- snats websitesnats.xyz
- Dataset list - A list of the biggest machine learning datasetsdatasetlist.com
- [2509.14786] Pre-training under infinite computearxiv.org
- Data Management For Large Language Models: A Surveyarxiv.org