Muon: An optimizer for hidden layers in neural networks | Keller Jordan blog
Muon is an optimizer for the hidden layers in neural networks. It is used in the current training speed records for both NanoGPT and CIFAR-10 speedrunning. Many empirical results using Muon have already been posted, so this writeup will focus mainly on Muon’s design. First we will define Muon and provide an overview of the empirical results it has achieved so far. Then we will discuss its design in full detail, including connections to prior research and our best understanding of why it works.
Muon is an optimizer for the hidden layers in neural networks. It is used in the current training speed records for both NanoGPT and CIFAR-10 speedrunning . Many empirical results using Muon have already been posted, so this writeup will focus mainly on Muon’s design. First we will define Muon and provide an overview of the empirical results it has achieved so far. Then we will discuss its design in full detail, including connections to prior research and our best understanding of why it works. Finally we will end with a discussion on standards of evidence in optimization research. Definition
saved by
- Jennifer Zhao
- Sarah Pan
- Yudhister Joel Kumar
- Nathan Chen
- Alyssa Yu
- Zeyneb Kaya
- Ishaan Panigrahi
- SUJASH AGRAWAL
related reading
- Deriving Muonjeremybernste.in
- Understanding Muonlakernewhouse.com
- Modular Manifolds - Thinking Machines Labthinkingmachines.ai
- Does Muon improve regulatory DNA learning? Part 1. — Origin Bioorigin.bio
- [2604.04891] Muon Dynamics as a Spectral Wasserstein Flowarxiv.org
- NL.pdfabehrouz.github.io
- Online KL Shampoo | Tildeblog.tilderesearch.com
- Muon Outperforms Adam in Tail-End Associative Memory Learningarxiv.org
- [2502.16982] Muon is Scalable for LLM Trainingarxiv.org
- The Practitioner's Guide to the Maximal Update Parameterization | EleutherAI Blogblog.eleuther.ai
- A Proof of Learning Rate Transfer under $\mu$Parxiv.org
- Notesandyrdt.com