The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale Guilherme Penedo Hynek Kydlíček Loubna Ben allal Anton Lozhkov Margaret Mitchell Colin Raffel Leandro Von Werra Thomas Wolf Hugging Face Abstract The performance of a large language model (LLM) depends heavily on the quality and size of its pretraining dataset. However, the pretraining datasets for state-of-the-art open LLMs like Llama 3 and Mixtral are not publicly available and very little is known about how they were created. In this work, we introduce FineWeb, a 15-trillion token dataset derived from 96 Common Crawl
Explore this link on the map →related reading
- DataRater: Meta-Learned Dataset Curationarxiv.org
- A Bitter Lesson for Data Filteringarxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretrainingdatologyai.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Dataset list - A list of the biggest machine learning datasetsdatasetlist.com
- Data Management For Large Language Models: A Surveyarxiv.org
- RedPajama, a project to create leading open-source models, starts by reproducing LLaMA training dataset of over 1.2 trillion tokenstogether.xyz
- snats websitesnats.xyz
- The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities (Version 1.0)arxiv.org