Synthetic Pretraining | Vintage Data
vintagedata.org · 4,034 words · saved by 1 readers
Old data, new models
Synthetic Pretraining | Vintage Data Synthetic pretraining Pierre-Carl Langlais, February 1, 2026 Pretraining data infrastructure used to be the most conservative part of a fast-moving AI world. Since GPT-3 we have been mostly scaling the usual mix of web crawls peppered with a few more select sources (including, controversially, digitized books). This is finally changing. In 2025, several major releases used extensive synthetic datasets before mid-training happens: Minimax , Trinity , K2 /K2.5, Nemotron-3 and, more speculatively, GPT-OSS . At Pleias we even experimented with full synthetic tr
related reading
- Training as we know it will end | Vintage Datavintagedata.org
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretrainingdatologyai.com
- How to Generate and Use Synthetic Data for Finetuningeugeneyan.com
- Autodata: An agentic data scientist to create high quality synthetic dataalphaxiv.org
- As Rocks May Think | Eric Jangevjang.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- [2409.19759] Balancing Cost and Effectiveness of Synthetic Data Generation Strategies for LLMsarxiv.org
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- Alignment pretraining could backfire — LessWronglesswrong.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com