Infinite Limits of Neural Networks - Kempner Institute
The performance of deep learning models improves with model size and dataset size in remarkably regular and predictable ways [1][2]. However, less is known about what kind of limiting behavior these models approach as their model size and dataset size approaches infinity. In this blog post, we aim to give the reader a relatively accessible introduction to various infinite parameter limits of neural networks. Beyond this, we aim to answer a pressing question of whether these theoretical limits actually translate to anything practically meaningful. Are large-scale language and vision models anywhere near these infinite limits? If not, in what ways do they differ? In studying very wide and deep networks, a crucial role will be played by the parameterization of the network. A parameterization is a rule for going from a given size network to a wider (or deeper) one. More technically, it is defined by how width and depth enter when one defines the initialization, forward pass, and gradient u
Deeper Learning Blog Infinite Limits of Neural Networks May 13, 2024 Part 1 of a two-part blog post covering recent findings from the authors By: Alex Atanasov, Blake Bordelon, and Cengiz Pehlevan Authors’ note: Papers discussed in this post include “Feature Learning Networks are Consistent Across Widths at Realistic Scales” coauthored by Nikhil Vyas (equal first author), Depen Morwani, and Sabarish Sainathan; and “Depthwise Hyperparameter Transfer: Dynamics and Scaling Limit” coauthored by Lorenzo Noci (equal first author), Mufan Bill Li, and Boris Hanin. The per
Explore this link on the map →saved by
related reading
- On neural scaling and the quanta hypothesisericjmichaud.com
- Some Math behind Neural Tangent Kernel | Lil'Loglilianweng.github.io
- Greg Yang | Professional pagethegregyang.com
- [2304.03408] Dynamics of Finite Width Kernel and Prediction Fluctuations in Mean Field Neural Networksarxiv.org
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- The Scaling Hypothesis · Gwern.netgwern.net
- Understanding the Neural Tangent Kernel – EigenTaleseigentales.com
- Understanding the Neural Tangent Kernel – EigenTalesrajatvd.github.io
- The Practitioner’s Guide to the Maximal Update Parameterization - Cerebrascerebras.ai
- The Little Book of Deep Learningfleuret.org
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com