[2605.06546] Efficient Pre-Training with Token Superposition
Abstract:Pre-training of Large Language Models is often prohibitively expensive and inefficient at scale, requiring complex and invasive modifications in order to achieve high data throughput. In this work, we present Token-Superposition Training (TST), a simple drop-in method that significantly improves the data throughput per FLOPs during pre-training without modifying the parallelism, optimizer, tokenizer, data, or model architecture. TST is done in two phases: (i) A highly efficient superposition phase where we combine many contiguous tokens into one bag and train using a multi-hot cross-entropy (MCE) objective, and (ii) a recovery phase where we revert back to standard training. We extensively evaluate TST on the scale of 270M and 600M parameters and validate on 3B and a 10B A1B mixture of experts model, demonstrating that it is highly robust in different settings. Ultimately, TST consistently outperforms baseline loss and downstream evaluations, and under equal-loss settings, TST yields up to a 2.5x reduction in total pre-training time at the 10B A1B scale.
# link_p85xjwxhxj.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Bowen Peng; Théo Gigant; Jeffrey Quesnelle - Creator=arXiv GenPDF (tex2pdf:a6404ea) - Custom.DOI=https://doi.org/10.48550/arXiv.2605.06546 - Custom.License=http://creativecommons.org/licenses/by/4.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.28 (TeX Live 2025) kpathsea version 6.4.1 - Custom.arXivID=https://arxiv.org/abs/2605.06546v2 - Producer=pikepdf 8.15.1 - Title=Efficient Pr
Explore this link on the map →saved by
related reading
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- The Bitter Lesson is coming for Tokenization – ⛰️ lucalplucalp.dev
- Composer2.pdfcursor.com
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- [2509.14786] Pre-training under infinite computearxiv.org
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- 2403.09611.pdfarxiv.org
- The Annotated Transformernlp.seas.harvard.edu
- RL is even more information inefficient than you thoughtdwarkesh.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- Large Language Diffusion Modelsarxiv.org