[2601.05034] How to Set the Batch Size for Large-Scale Pre-training?
Abstract:The concept of Critical Batch Size, as pioneered by OpenAI, has long served as a foundational principle for large-scale pre-training. However, with the paradigm shift towards the Warmup-Stable-Decay (WSD) learning rate scheduler, we observe that the original theoretical framework and its underlying mechanisms fail to align with new pre-training dynamics. To bridge this gap between theory and practice, this paper derives a revised E(S) relationship tailored for WSD scheduler, characterizing the trade-off between training data consumption E and steps S during pre-training. Our theoretical analysis reveals two fundamental properties of WSD-based pre-training: 1) B_min, the minimum batch size threshold required to achieve a target loss, and 2) B_opt, the optimal batch size that maximizes data efficiency by minimizing total tokens. Building upon these properties, we propose a dynamic Batch Size Scheduler. Extensive experiments demonstrate that our revised formula precisely captures the dynamics of large-scale pre-training, and the resulting scheduling strategy significantly enhances both training efficiency and final model quality.
[2601.05034] How to Set the Batch Size for Large-Scale Pre-training? Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Artificial Intelligence arXiv:2601.05034 (cs) [Submitted on 8 Jan 2026 ( v1 ), last revised 9 Jan 2026 (this version, v2)] Title: How to Set the Batch Size for Large-Scale Pre-training? Authors: Yunhua Zhou , Junhao Huang , Shuhao Xing , Yechen Zhang , Runyu Peng , Qiping Guo , Xipeng Qiu View a PDF of the paper titled How to Set the Batch Size for Large-Scale Pre-tr
Explore this link on the map →related reading
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- [2509.14786] Pre-training under infinite computearxiv.org
- [2503.04715] Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretrainingarxiv.org
- The Practitioner’s Guide to the Maximal Update Parameterization - Cerebrascerebras.ai
- [1706.02677] Accurate, Large Minibatch SGD: Training ImageNet in 1 Hourarxiv-vanity.com
- [2603.21191] On the Role of Batch Size in Stochastic Conditional Gradient Methodsarxiv.org
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Jane Street Blog - Does batch size matter?blog.janestreet.com
- [2605.06546] Efficient Pre-Training with Token Superpositionarxiv.org
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- Modern Pretraining Strategies: A Hands-On Guidetheneuralmaze.substack.com