Greg Yang | Professional page
I am currently developing a framework called Tensor Programs for understanding large neural networks. This is the theoretical foundation that gave rise to the hyperparameter transfer paradigm for tuning enormous neural networks like GPT-3. Ultimately, my goal is a Theory of Everything for large scale deep learning that 1) tells us the optimal way of scaling neural networks and 2) provides robust understanding to such models so as to guide safety and alignment efforts. Informally, a TP is just a composition of matrix multiplication and coordinatewise nonlinearities. It turns out that practically any computation in Deep Learning can be expressed as a TP, e.g. training a transformer on wikipedia data. Simultaneously, any Tensor Program has an “infinite-width” limit which can be derived from the program itself (through what’s called the Master Theorem). So this gives a universal way of taking the infinite-width limit of any deep learning computation, e.g. training an infinite-width transfo
Greg Yang | Professional page About Me I am a mathematician at xAI . Previously I was a researcher at Microsoft Research . I am currently developing a framework called Tensor Programs for understanding large neural networks. This is the theoretical foundation that gave rise to the hyperparameter transfer paradigm for tuning enormous neural networks like GPT-3. Ultimately, my goal is a Theory of Everything for large scale deep learning that 1) tells us the optimal way of scaling neural networks and 2) provides robust understanding to such models so as to guide safety and alignment efforts. Tens
saved by
related reading
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- Infinite Limits of Neural Networks - Kempner Institutekempnerinstitute.harvard.edu
- [2605.01172] A Theory of Generalization in Deep Learningarxiv.org
- GitHub - adam-maj/deep-learning: A deep-dive on the entire history of deep-learninggithub.com
- The Scaling Hypothesis · Gwern.netgwern.net
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- Some Math behind Neural Tangent Kernel | Lil'Loglilianweng.github.io
- A Proof of Learning Rate Transfer under $\mu$Parxiv.org
- The Practitioner’s Guide to the Maximal Update Parameterization - Cerebrascerebras.ai
- nn-notes.pdfboris-hanin.github.io
- How To Scale Your Modeljax-ml.github.io
- The Practitioner's Guide to the Maximal Update Parameterization | EleutherAI Blogblog.eleuther.ai