[2601.08393] Controlled LLM Training on Spectral Sphere
Abstract:Scaling large models requires optimization strategies that ensure rapid convergence grounded in stability. Maximal Update Parametrization ($\boldsymbol{\mu}$P) provides a theoretical safeguard for width-invariant $\Theta(1)$ activation control, whereas emerging optimizers like Muon are only ``half-aligned'' with these constraints: they control updates but allow weights to drift. To address this limitation, we introduce the \textbf{Spectral Sphere Optimizer (SSO)}, which enforces strict module-wise spectral constraints on both weights and their updates. By deriving the steepest descent direction on the spectral sphere, SSO realizes a fully $\boldsymbol{\mu}$P-aligned optimization process. To enable large-scale training, we implement SSO as an efficient parallel algorithm within Megatron. Through extensive pretraining on diverse architectures, including Dense 1.7B, MoE 8B-A1B, and 200-layer DeepNet models, SSO consistently outperforms AdamW and Muon. Furthermore, we observe significant practical stability benefits, including improved MoE router load balancing, suppressed outliers, and strictly bounded activations.
Controlled LLM Training on Spectral Sphere * Tian Xie1 Haoming Luo2 Haoyu Tang2 Yiwen Hu2 Jason Klein Liu4 Qingnan Ren1 Yang Wang1 Wayne Xin Zhao2,4 Rui Yan3 Bing Su2 Chong Luo1…
saved by
related reading
- The Practitioner’s Guide to the Maximal Update Parameterization - Cerebrascerebras.ai
- The Practitioner's Guide to the Maximal Update Parameterization | EleutherAI Blogblog.eleuther.ai
- Deriving Muonjeremybernste.in
- Muon: An optimizer for hidden layers in neural networks | Keller Jordan blogkellerjordan.github.io
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Modular Manifolds - Thinking Machines Labthinkingmachines.ai
- Understanding Muonlakernewhouse.com
- How To Scale Your Modeljax-ml.github.io
- A Spectral Condition for Feature Learningarxiv.org
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- [2502.16982] Muon is Scalable for LLM Trainingarxiv.org
- Does Muon improve regulatory DNA learning? Part 1. — Origin Bioorigin.bio