✳flâneur — a map of the web's best reading
The Practitioner’s Guide to the Maximal Update Parameterization - Cerebras
cerebras.ai · 4,907 words · saved by 2 readers
Cerebras is the go-to platform for fast and effortless AI training. Learn more at cerebras.ai.
The Practitioner’s Guide to the Maximal Update Parameterization - Cerebras Skip to main content Sep 23 2024 The Practitioner’s Guide to the Maximal Update Parameterization - Cerebras Nolan Dey Quentin Anthony Joel Hestness Introduction Maximal Update Parameterization (µP) offers significant advantages for neural network training, but its adoption has been limited due to the complexity of the underlying math and the challenges in implementation. This guide aims to lower those barriers by providing a clear and practical overview of µP. By using µP, you can achieve stable hyperparameters across m
Explore this link on the map →saved by
related reading
- The Practitioner's Guide to the Maximal Update Parameterization | EleutherAI Blogblog.eleuther.ai
- Does Muon improve regulatory DNA learning? Part 1. — Origin Bioorigin.bio
- Deriving Muonjeremybernste.in
- How To Scale Your Modeljax-ml.github.io
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Greg Yang | Professional pagethegregyang.com
- The Decade of Deep Learning | Leo Gaobmk.sh
- NL.pdfabehrouz.github.io
- Infinite Limits of Neural Networks - Kempner Institutekempnerinstitute.harvard.edu
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- The Scaling Hypothesis · Gwern.netgwern.net
- arxiv.org/pdf/2512.24880#page=3.56arxiv.org