✳flâneur — a map of the web's best reading
The Unreasonable Effectiveness of Recurrent Neural Networks
karpathy.github.io · 8,155 words · saved by 1 readers
Musings of a Computer Scientist.
There’s something magical about Recurrent Neural Networks (RNNs). I still remember when I trained my first recurrent network for Image Captioning . Within a few dozen minutes of training my first baby model (with rather arbitrarily-chosen hyperparameters) started to generate very nice looking descriptions of images that were on the edge of making sense. Sometimes the ratio of how simple your model is to the quality of the results you get out of it blows past your expectations, and this was one of those times. What made this result so shocking at the time was that the common wisdom was that RNN
Explore this link on the map →related reading
- The Unreasonable Effectiveness of Recurrent Neural Networkskarpathy.github.io
- Recurrent Neural Networks Tutorial, Part 1 – Introduction to RNNs · Denny's Blogdennybritz.com
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- Ilya 30u30arc.net
- Recurrent neural network - Wikipediaen.wikipedia.org
- Understanding LSTM Networks -- colah's blogcolah.github.io
- Attention and Augmented Recurrent Neural Networksdistill.pub
- GitHub - karpathy/char-rnn: Multi-layer Recurrent Neural Networks (LSTM, GRU, RNN) for character-level language models in Torch · GitHubgithub.com
- [2606.06479] Pretraining Recurrent Networks without Recurrencearxiv.org
- An Overview of Deep Learning for Curious People | Lil'Loglilianweng.github.io
- What Is ChatGPT Doing … and Why Does It Work?-Stephen Wolfram Writingswritings.stephenwolfram.com
- A History of Large Language Modelsgregorygundersen.com