[2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraints
Abstract:As language models scale, the amount of data they require grows -- yet many target data sources, such as low-resource languages or specialized domains, are inherently limited in size. A common strategy is to mix this scarce but valuable target data with abundant generic data, which presents a fundamental trade-off: too little target data in the mixture underexposes the model to the target domain, while too much target data repeats the same examples excessively, yielding diminishing returns and eventual overfitting. We study this trade-off across more than 2,000 language-model training runs spanning multiple model and target dataset sizes, as well as several data types, including multilingual, domain-specific, and quality-filtered mixtures. Across all settings, we find that repetition is a central driver of target-domain performance, and that mixture training tolerates much higher repetition than single-source training: scarce target corpora can be reused 15-20 times, with the optimal number of repetitions depending on the target data size, compute budget, and model scale. Next, we introduce a repetition-aware mixture scaling law that accounts for the decreasing value of repeated target tokens and the regularizing role of generic data. Optimizing the scaling law provides a principled way to compute effective mixture configurations, yielding practical mixture recommendations for pretraining under data constraints.
# link_1d3groarnfp.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Anastasiia Sedova; Skyler Seto; Natalie Schluter; Pierre Ablin - Creator=arXiv GenPDF (tex2pdf:a6404ea) - Custom.DOI=https://doi.org/10.48550/arXiv.2605.12715 - Custom.License=http://arxiv.org/licenses/nonexclusive-distrib/1.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.28 (TeX Live 2025) kpathsea version 6.4.1 - Custom.arXivID=https://arxiv.org/abs/2605.12715v2 - Producer=pikepd
Explore this link on the map →saved by
related reading
- [2605.15220] Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Timearxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- A Bitter Lesson for Data Filteringarxiv.org
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- On neural scaling and the quanta hypothesisericjmichaud.com
- chinchilla's wild implications — AI Alignment Forumalignmentforum.org
- Data Management For Large Language Models: A Surveyarxiv.org
- [2509.14786] Pre-training under infinite computearxiv.org
- chinchilla's wild implications — LessWronglesswrong.com
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io