BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
In this post we present BeyondWeb, the synthetic data component of our pretraining data curation pipeline. BeyondWeb is a synthetic data generation framework that leverages targeted document rephrasing to yield diverse, relevant, and information-dense synthetic pretraining data. Here we show that BeyondWeb substantially outperforms existing state-of-the-art public pretraining datasets and share some of the lessons we learned along the way about why it's so hard to generate high-quality synthetic pretraining data. To learn more about our full data curation pipeline beyond BeyondWeb, see this blog post. Motivation Recent advances in LLM pretraining have shown that simply scaling data quantity eventually leads to diminishing returns, hitting a data wall, beyond which data of high information-density is prohibitively scarce. Synthetic data has emerged as a powerful complement to scarce, high-quality web text. The success of synthetic data methods has led research organizations to pour subs
Research Updates 18 Aug, 2025 BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining Written by D DatologyAI Published on 18 Aug, 2025 Share on Table of Contents Executive Summary 1. Introduction 2. A Tale of Two Approaches for Synthetic Pretraining Data Generation 3. Introducing BeyondWeb 4. Systematically Evaluating Synthetic Data 5. Future Directions 6. Conclusion Get in Touch! Contributions and Acknowledgements *NOTE : If you prefer a PDF, the ArXiv version of this post can be found here . Executive Summary In this post we present BeyondWeb , the synthetic data compo
Explore this link on the map →saved by
related reading
- Synthetic Pretraining | Vintage Datavintagedata.org
- How to Generate and Use Synthetic Data for Finetuningeugeneyan.com
- Training as we know it will end | Vintage Datavintagedata.org
- The Only Important Technology Is The Internet - Kevin Lukevinlu.ai
- Machine Learning for Synthetic Data Generation: A Reviewarxiv.org
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraintsarxiv.org
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scalearxiv.org
- Data Management For Large Language Models: A Surveyarxiv.org
- [2503.18866] Reasoning to Learn from Latent Thoughtsarxiv.org
- Large Language Diffusion Modelsarxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai