✳flâneur — a map of the web's best reading
The Practitioner's Guide to the Maximal Update Parameterization | EleutherAI Blog
blog.eleuther.ai · 5,057 words · saved by 3 readers
Exploring the implementation details of muTransfer
Table of Contents Introduction Why you should use μP 1. Stable optimum HPs across scale (μTransfer) 2. Improved loss at large scale due to improved HP tuning 3. Stable training - significantly decreased danger of instability at large scale 4. More predictable scaling due to μTransfer μP enables better research A Simple Approach to the μP Math Basic Building Block: Controlled Activation Magnitudes Operations in a training step Practitioner's guide to μP Implementation Coordinate check test μTransfer test Transferring optimal HPs from a small scale to a large scale Conclusion Citation Footnotes
Explore this link on the map →saved by
related reading
- The Practitioner’s Guide to the Maximal Update Parameterization - Cerebrascerebras.ai
- Deriving Muonjeremybernste.in
- Does Muon improve regulatory DNA learning? Part 1. — Origin Bioorigin.bio
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- On neural scaling and the quanta hypothesisericjmichaud.com
- Marketplace: my first attempt at training without backprop on GPU efficiently – Fang-Pen's coding notefangpenlin.com
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- How To Scale Your Modeljax-ml.github.io
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Greg Yang | Professional pagethegregyang.com
- Infinite Limits of Neural Networks - Kempner Institutekempnerinstitute.harvard.edu
- arxiv.org/pdf/2512.24880#page=3.56arxiv.org