A Proof of Learning Rate Transfer under $\mu$P
arxiv.org · 6,241 words · saved by 2 readers
N/A
A Proof of Learning Rate Transfer under µP Soufiane Hayou Department of Applied Mathematics and Statistics Johns Hopkins University arXiv:2511.01734v3 [stat.ML] 24 Feb 2026 Abstract understand training dynamics of large-width neural…
saved by
related reading
- The Practitioner’s Guide to the Maximal Update Parameterization - Cerebrascerebras.ai
- The Practitioner's Guide to the Maximal Update Parameterization | EleutherAI Blogblog.eleuther.ai
- Deriving Muonjeremybernste.in
- Infinite Limits of Neural Networks - Kempner Institutekempnerinstitute.harvard.edu
- Greg Yang | Professional pagethegregyang.com
- NL.pdfabehrouz.github.io
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- nn-notes.pdfboris-hanin.github.io
- Some Math behind Neural Tangent Kernel | Lil'Loglilianweng.github.io
- Muon: An optimizer for hidden layers in neural networks | Keller Jordan blogkellerjordan.github.io
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- A Spectral Condition for Feature Learningarxiv.org