flâneur — a map of the web's best reading

The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

arxiv.org · 11,749 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.

The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale Guilherme Penedo Hynek Kydlíček Loubna Ben allal Anton Lozhkov Margaret Mitchell Colin Raffel Leandro Von Werra Thomas Wolf Hugging Face Abstract The performance of a large language model (LLM) depends heavily on the quality and size of its pretraining dataset. However, the pretraining datasets for state-of-the-art open LLMs like Llama 3 and Mixtral are not publicly available and very little is known about how they were created. In this work, we introduce FineWeb, a 15-trillion token dataset derived from 96 Common Crawl

Explore this link on the map →

related reading