✳flâneur — a map of the web's best reading
A Recipe for Training Neural Networks
karpathy.github.io · 3,882 words · saved by 13 readers
Musings of a Computer Scientist.
Some few weeks ago I posted a tweet on “the most common neural net mistakes”, listing a few common gotchas related to training neural nets. The tweet got quite a bit more engagement than I anticipated (including a webinar :)). Clearly, a lot of people have personally encountered the large gap between “here is how a convolutional layer works” and “our convnet achieves state of the art results”. So I thought it could be fun to brush off my dusty blog to expand my tweet to the long form that this topic deserves. However, instead of going into an enumeration of more common errors or fleshing them
Explore this link on the map →saved by
- Winnie Xu
- Claire Wang
- Ivy Zhang
- Jordan Sucher
- Andrew Kachnic
- Kenson Hui
- Jacob G-W
- Gloria Ma
- Tomi Jaga
- Liam Hinzman
- Curtis Chong
- Akira Yoshiyama
related reading
- A Recipe for Training Neural Networkskarpathy.github.io
- Bayesian Neural Networkscs.toronto.edu
- The Decade of Deep Learning | Leo Gaobmk.sh
- Neural network training makes beautiful fractals | Jascha’s blogsohl-dickstein.github.io
- Neural Networks, Manifolds, and Topology -- colah's blogcolah.github.io
- The Unreasonable Effectiveness of Recurrent Neural Networkskarpathy.github.io
- Feature Visualizationdistill.pub
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- Neural networks and deep learningneuralnetworksanddeeplearning.com
- Feature Visualizationdistill.pub
- The Little Book of Deep Learningfleuret.org
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org