[2607.16051] Loop the Loopies!
Abstract:We present the Loopie series, consisting of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6B-parameter model with 0.6B active parameters. Looped Transformers have long faced a challenge: given an N times increase in pre-training compute, increasing the parameter count by a factor of N usually outperforms looping a model N times. Loopie addresses this challenge. Extensive ablation studies, including comparisons with a vanilla 30B-A3B model, show that Loopie substantially outperforms vanilla Transformer baselines trained with the same compute budget. With a novel post-training method, Loopie develops strong reasoning abilities and achieves frontier-level reasoning performance.
Loop the Loopies! Zitian Gao Yilong Chen Yihao Xiao Xinyu Yang Ran Tao Joey Zhou Bryan Dai IQuest Research See the full author contributions here. Models: Loopie-20B-A2B…
saved by
related reading
- DeepSeek-V3: A Large-Scale MoE Pretraining Benchmark for MLPerf Training v6.0mlcommons.org
- [2509.14786] Pre-training under infinite computearxiv.org
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Training-Free Looped Transformersarxiv.org
- GitHub - huskydoge/Awesome-Loop-Models: A curated list of papers and selected technical blogs on Loop Models.github.com
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- [2101.03961] Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsityarxiv.org
- [2607.13491] DeepLoop: Depth Scaling for Looped Transformersarxiv.org
- Composer2.pdfcursor.com
- How We Build Trillion Parameter Reasoning RL with 10% GPUsmacaron.im
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Papers I’ve read this week, Mixture of Experts editionfinbarrtimbers.substack.com