learning theory
people.eecs.berkeley.edu · 2,800 words · saved by 1 readers
N/A
Residual Networks; Batch Normalization; AdamW 141 23 Residual Networks; Batch Normalization; AdamW TRAINING DEEP NETWORKS Most influential ideas: ResNets, batch normalization, layer normalization. [These ideas enable deep networks to train. They reduce the likelihood of encountering the vanishing gradi- ent or exploding gradient problems, but don’t quite eliminate them.] Batch Normalization [Batch normalization has played a huge role in making it easier to train very deep neural networks since its introduction in 2015, and it’s still a mainstay…
related reading
- The Decade of Deep Learning | Leo Gaobmk.sh
- Residual neural network - Wikipediaen.wikipedia.org
- 1512.03385arxiv.org
- bachlechner21a.pdfproceedings.mlr.press
- A Recipe for Training Neural Networkskarpathy.github.io
- ResNets: Why do they perform better than Classic ConvNets? (Conceptual Analysis) | Towards Data Sciencetowardsdatascience.com
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- Batch Normalization in Convolutional Neural Networks | Baeldung on Computer Sciencebaeldung.com
- What is Residual Connection? | Towards Data Sciencetowardsdatascience.com
- [1607.06450] Layer Normalizationarxiv.org
- Convolutional Neural Networks, Explained | Towards Data Sciencetowardsdatascience.com
- The Little Book of Deep Learningfleuret.org