Nonlinear Computation in Deep Linear Networks
We've shown that deep linear networks — as implemented using floating-point arithmetic — are not actually linear and can perform nonlinear computation. We used evolution strategies to find parameters in linear networks that exploit this trait, letting us solve non-trivial problems. Neural networks consist of stacks of a linear layer followed by
September 29, 2017 4 minute read We've shown that deep linear networks — as implemented using floating-point arithmetic — are not actually linear and can perform nonlinear computation. We used evolution strategies to find parameters in linear networks that exploit this trait, letting us solve non-trivial problems. Neural networks consist of stacks of a linear layer followed by a nonlinearity like tanh or rectified linear unit. Without the nonlinearity, consecutive linear layers would be in theory mathematically equivalent to a single linear layer. So it's a surprise that floating point arithme
Explore this link on the map →related reading
- The Little Book of Deep Learningfleuret.org
- Neural networks and deep learningneuralnetworksanddeeplearning.com
- What Would Non-Linear Features Actually Look Like? — Liv Gortonlivgorton.com
- Making Deep Learning go Brrrr From First Principleshorace.io
- Neural Networks, Types, and Functional Programming -- colah's blogcolah.github.io
- The Decade of Deep Learning | Leo Gaobmk.sh
- Yes you should understand backprop | by Andrej Karpathy | Mediumkarpathy.medium.com
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- Neural network training makes beautiful fractals | Jascha’s blogsohl-dickstein.github.io
- 2404.17625arxiv.org
- Latest | Epoch AIepochai.org
- Estimating training compute of deep learning models | Epoch AIepochai.org