[1802.05296] Stronger generalization bounds for deep nets via a compression approach
Abstract:Deep nets generalize well despite having more parameters than the number of training samples. Recent works try to give an explanation using PAC-Bayes and Margin-based analyses, but do not as yet result in sample complexity bounds better than naive parameter counting. The current paper shows generalization bounds that're orders of magnitude better in practice. These rely upon new succinct reparametrizations of the trained net --- a compression that is explicit and efficient. These yield generalization bounds via a simple compression-based framework introduced here. Our results also provide some theoretical justification for widespread empirical success in compressing deep nets. Analysis of correctness of our compression relies upon some newly identified \textquotedblleft noise stability\textquotedblright properties of trained deep nets, which are also experimentally verified. The study of these properties and resulting generalization bounds are also extended to convolutional nets, which had eluded earlier attempts on proving generalization.
Abstract:Deep nets generalize well despite having more parameters than the number of training samples. Recent works try to give an explanation using PAC-Bayes and Margin-based analyses, but do not as yet result in sample complexity bounds better than naive parameter counting. The current paper shows generalization bounds that're orders of magnitude better in practice. These rely upon new succinct reparametrizations of the trained net --- a compression that is explicit and efficient. These yield generalization bounds via a simple compression-based framework introduced here. Our results also pro
Explore this link on the map →related reading
- arxiv.org/pdf/1805.08522arxiv.org
- [2503.02113] Deep Learning is Not So Mysterious or Differentarxiv.org
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- [1805.08522] Deep learning generalizes because the parameter-function map is biased towards simple functionsarxiv.org
- The Scaling Hypothesis · Gwern.netgwern.net
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- The Decade of Deep Learning | Leo Gaobmk.sh
- The Little Book of Deep Learningfleuret.org
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- [2605.01172] A Theory of Generalization in Deep Learningarxiv.org
- Just Ask for Generalization | Eric Jangevjang.com
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com