Greg Yang | Professional page
I am currently developing a framework called Tensor Programs for understanding large neural networks. This is the theoretical foundation that gave rise to the hyperparameter transfer paradigm for tuning enormous neural networks like GPT-3. Ultimately, my goal is a Theory of Everything for large scale deep learning that 1) tells us the optimal way of scaling neural networks and 2) provides robust understanding to such models so as to guide safety and alignment efforts. Informally, a TP is just a composition of matrix multiplication and coordinatewise nonlinearities. It turns out that practically any computation in Deep Learning can be expressed as a TP, e.g. training a transformer on wikipedia data. Simultaneously, any Tensor Program has an “infinite-width” limit which can be derived from the program itself (through what’s called the Master Theorem). So this gives a universal way of taking the infinite-width limit of any deep learning computation, e.g. training an infinite-width transfo
Greg Yang | Professional page About Me I am a mathematician at xAI . Previously I was a researcher at Microsoft Research . I am currently developing a framework called Tensor Programs for understanding large neural networks. This is the theoretical foundation that gave rise to the hyperparameter transfer paradigm for tuning enormous neural networks like GPT-3. Ultimately, my goal is a Theory of Everything for large scale deep learning that 1) tells us the optimal way of scaling neural networks and 2) provides robust understanding to such models so as to guide safety and alignment efforts. Tens
Explore this link on the map →saved by
related reading
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- Infinite Limits of Neural Networks - Kempner Institutekempnerinstitute.harvard.edu
- [2605.01172] A Theory of Generalization in Deep Learningarxiv.org
- The Scaling Hypothesis · Gwern.netgwern.net
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- Some Math behind Neural Tangent Kernel | Lil'Loglilianweng.github.io
- The Practitioner’s Guide to the Maximal Update Parameterization - Cerebrascerebras.ai
- How To Scale Your Modeljax-ml.github.io
- On neural scaling and the quanta hypothesisericjmichaud.com
- The Practitioner's Guide to the Maximal Update Parameterization | EleutherAI Blogblog.eleuther.ai
- Understanding the Neural Tangent Kernel – EigenTaleseigentales.com
- The Little Book of Deep Learningfleuret.org