The Practitioner’s Guide to the Maximal Update Parameterization - Cerebras
cerebras.ai · 4,907 words · saved by 2 readers
Cerebras is the go-to platform for fast and effortless AI training. Learn more at cerebras.ai.
The Practitioner’s Guide to the Maximal Update Parameterization - Cerebras Skip to main content Sep 23 2024 The Practitioner’s Guide to the Maximal Update Parameterization - Cerebras Nolan Dey Quentin Anthony Joel Hestness Introduction Maximal Update Parameterization (µP) offers significant advantages for neural network training, but its adoption has been limited due to the complexity of the underlying math and the challenges in implementation. This guide aims to lower those barriers by providing a clear and practical overview of µP. By using µP, you can achieve stable hyperparameters across m
saved by
related reading
- The Practitioner's Guide to the Maximal Update Parameterization | EleutherAI Blogblog.eleuther.ai
- A Proof of Learning Rate Transfer under $\mu$Parxiv.org
- Does Muon improve regulatory DNA learning? Part 1. — Origin Bioorigin.bio
- Deriving Muonjeremybernste.in
- How To Scale Your Modeljax-ml.github.io
- [2601.08393] Controlled LLM Training on Spectral Spherearxiv.org
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Greg Yang | Professional pagethegregyang.com
- Infinite Limits of Neural Networks - Kempner Institutekempnerinstitute.harvard.edu
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- A Spectral Condition for Feature Learningarxiv.org
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io