flâneur

learning theory

people.eecs.berkeley.edu · 2,800 words · saved by 1 readers

N/A

Residual Networks; Batch Normalization; AdamW 141 23 Residual Networks; Batch Normalization; AdamW TRAINING DEEP NETWORKS Most influential ideas: ResNets, batch normalization, layer normalization. [These ideas enable deep networks to train. They reduce the likelihood of encountering the vanishing gradi- ent or exploding gradient problems, but don’t quite eliminate them.] Batch Normalization [Batch normalization has played a huge role in making it easier to train very deep neural networks since its introduction in 2015, and it’s still a mainstay…

related reading