A Recipe for Training Neural Networks
karpathy.github.io · 3,882 words · saved by 15 readers
Musings of a Computer Scientist.
Some few weeks ago I posted a tweet on “the most common neural net mistakes”, listing a few common gotchas related to training neural nets. The tweet got quite a bit more engagement than I anticipated (including a webinar :)). Clearly, a lot of people have personally encountered the large gap between “here is how a convolutional layer works” and “our convnet achieves state of the art results”. So I thought it could be fun to brush off my dusty blog to expand my tweet to the long form that this topic deserves. However, instead of going into an enumeration of more common errors or fleshing them
saved by
- Winnie Xu
- Claire Wang
- Ivy Zhang
- Jordan Sucher
- Uzay Girit
- Andrew Kachnic
- Kenson Hui
- Jacob G-W
- Gloria Ma
- Tomi Jaga
- Liam Hinzman
- Curtis Chong
related reading
- A Recipe for Training Neural Networkskarpathy.github.io
- Bayesian Neural Networkscs.toronto.edu
- Neural network training makes beautiful fractals | Jascha’s blogsohl-dickstein.github.io
- Neural Networks, Manifolds, and Topology -- colah's blogcolah.github.io
- The Decade of Deep Learning | Leo Gaobmk.sh
- Neural networks and deep learningneuralnetworksanddeeplearning.com
- Feature Visualizationdistill.pub
- nn-notes.pdfboris-hanin.github.io
- Feature Visualizationdistill.pub
- CS231n Deep Learning for Computer Visioncs231n.github.io
- The Little Book of Deep Learningfleuret.org
- How our data shaped neural architecture discovery, and how automation can reshape the future | Core Automationcoreauto.com