flâneur — a map of the web's best reading

The Generalization Mystery: Sharp vs Flat Minima

inference.vc · 1,913 words · saved by 1 readers

Inevitably, I started thinking more generally about flat and sharp minima and generalization, so rather than describing these papers in details, I ended up dumping some thoughts of my own. Feedback and pointers to literature are welcome, as always The loss surface of deep nets tends to have many local minima. Many of these might be equally good in terms of training error, but they may have widely different generalization performance, i.e. an network with minimal training loss might perform very well, or very poorly on a held-out training set. Interestingly, stochastic gradient descent (SGD) with small batchsizes appears to locate minima with better generalization properties than large-batch SGD. So the big question is: what measurable property of a local minimum can we use to predict generalization properties? And how does this relate to SGD? There is speculation dating back to at least Hochreiter and Schmidhuber (1997) that the flatness of the minimum is a good measure to look at. How

January 18, 2018 The Generalization Mystery: Sharp vs Flat Minima I set out to write about the following paper I saw people talk about on twitter and reddit: Hao Li, Zheng Xu, Gavin Taylor, Tom Goldstein Visualizing the Loss Landscape of Neural Nets It's related to this pretty insightful paper: Laurent Dinh, Razvan Pascanu, Samy Bengio, Yoshua Bengio (2017) Sharp Minima Can Generalize For Deep Nets Inevitably, I started thinking more generally about flat and sharp minima and generalization, so rather than describing these papers in details, I ended up dumping some thoughts of my own. Feedback

Explore this link on the map →

related reading