The Practitioner's Guide to the Maximal Update Parameterization | EleutherAI Blog
blog.eleuther.ai · 5,057 words · saved by 3 readers
Exploring the implementation details of muTransfer
Table of Contents Introduction Why you should use μP 1. Stable optimum HPs across scale (μTransfer) 2. Improved loss at large scale due to improved HP tuning 3. Stable training - significantly decreased danger of instability at large scale 4. More predictable scaling due to μTransfer μP enables better research A Simple Approach to the μP Math Basic Building Block: Controlled Activation Magnitudes Operations in a training step Practitioner's guide to μP Implementation Coordinate check test μTransfer test Transferring optimal HPs from a small scale to a large scale Conclusion Citation Footnotes
saved by
related reading
- The Practitioner’s Guide to the Maximal Update Parameterization - Cerebrascerebras.ai
- Deriving Muonjeremybernste.in
- Does Muon improve regulatory DNA learning? Part 1. — Origin Bioorigin.bio
- A Proof of Learning Rate Transfer under $\mu$Parxiv.org
- [2601.08393] Controlled LLM Training on Spectral Spherearxiv.org
- [2608.13335] Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Lawsarxiv.org
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- How To Scale Your Modeljax-ml.github.io
- Muon: An optimizer for hidden layers in neural networks | Keller Jordan blogkellerjordan.github.io
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Greg Yang | Professional pagethegregyang.com