[1609.04836] On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Abstract:The stochastic gradient descent (SGD) method and its variants are algorithms of choice for many Deep Learning tasks. These methods operate in a small-batch regime wherein a fraction of the training data, say $32$-$512$ data points, is sampled to compute an approximation to the gradient. It has been observed in practice that when using a larger batch there is a degradation in the quality of the model, as measured by its ability to generalize. We investigate the cause for this generalization drop in the large-batch regime and present numerical evidence that supports the view that large-batch methods tend to converge to sharp minimizers of the training and testing functions - and as is well known, sharp minima lead to poorer generalization. In contrast, small-batch methods consistently converge to flat minimizers, and our experiments support a commonly held view that this is due to the inherent noise in the gradient estimation. We discuss several strategies to attempt to help large-batch methods eliminate this generalization gap.
[1609.04836] On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Machine Learning arXiv:1609.04836 (cs) [Submitted on 15 Sep 2016 ( v1 ), last revised 9 Feb 2017 (this version, v2)] Title: On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima Authors: Nitish Shirish Keskar , Dheevatsa Mudigere , Jorge Nocedal , Mikhail Smelyanskiy , Ping Tak Peter Tang View a PDF of the paper titled
Explore this link on the map →related reading
- [2101.12176] On the Origin of Implicit Regularization in Stochastic Gradient Descentarxiv.org
- The Generalization Mystery: Sharp vs Flat Minimainference.vc
- Notes on the Origin of Implicit Regularization in SGDinference.vc
- The Little Book of Deep Learningfleuret.org
- Just Ask for Generalization | Eric Jangevjang.com
- [2605.01172] A Theory of Generalization in Deep Learningarxiv.org
- The Scaling Hypothesis · Gwern.netgwern.net
- [1706.02677] Accurate, Large Minibatch SGD: Training ImageNet in 1 Hourarxiv-vanity.com
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- Thoughts on Loss Landscapes and why Deep Learning works — LessWronglesswrong.com