flâneur — a map of the web's best reading

Synthetic Pretraining | Vintage Data

vintagedata.org · 4,034 words · saved by 1 readers

Old data, new models

Synthetic Pretraining | Vintage Data Synthetic pretraining Pierre-Carl Langlais, February 1, 2026 Pretraining data infrastructure used to be the most conservative part of a fast-moving AI world. Since GPT-3 we have been mostly scaling the usual mix of web crawls peppered with a few more select sources (including, controversially, digitized books). This is finally changing. In 2025, several major releases used extensive synthetic datasets before mid-training happens: Minimax , Trinity , K2 /K2.5, Nemotron-3 and, more speculatively, GPT-OSS . At Pleias we even experimented with full synthetic tr

Explore this link on the map →

related reading