[2503.02113] Deep Learning is Not So Mysterious or Different
Abstract:Deep neural networks are often seen as different from other model classes by defying conventional notions of generalization. Popular examples of anomalous generalization behaviour include benign overfitting, double descent, and the success of overparametrization. We argue that these phenomena are not distinct to neural networks, or particularly mysterious. Moreover, this generalization behaviour can be intuitively understood, and rigorously characterized, using long-standing generalization frameworks such as PAC-Bayes and countable hypothesis bounds. We present soft inductive biases as a key unifying principle in explaining these phenomena: rather than restricting the hypothesis space to avoid overfitting, embrace a flexible hypothesis space, with a soft preference for simpler solutions that are consistent with the data. This principle can be encoded in many model classes, and thus deep learning is not as mysterious or different from other model classes as it might seem. However, we also highlight how deep learning is relatively distinct in other ways, such as its ability for representation learning, phenomena such as mode connectivity, and its relative universality.
# link_13kqqiiyw0p.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Andrew Gordon Wilson - Creator=arXiv GenPDF (tex2pdf:) - Custom.DOI=https://doi.org/10.48550/arXiv.2503.02113 - Custom.License=http://arxiv.org/licenses/nonexclusive-distrib/1.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.25 (TeX Live 2023) kpathsea version 6.3.5 - Custom.arXivID=https://arxiv.org/abs/2503.02113v2 - Producer=pikepdf 8.15.1 - Title=Deep Learning is Not So Mysterio
Explore this link on the map →saved by
related reading
- Deep learning as program synthesis — LessWronglesswrong.com
- The paper that killed deep learning theory — LessWronglesswrong.com
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- arxiv.org/pdf/1805.08522arxiv.org
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- The Scaling Hypothesis · Gwern.netgwern.net
- [1805.08522] Deep learning generalizes because the parameter-function map is biased towards simple functionsarxiv.org
- Clare Lyle | What's grokking good for?clarelyle.com
- 4 – The Overfitting Iceberg – Machine Learning Blog | ML@CMU | Carnegie Mellon Universityblog.ml.cmu.edu
- Just Ask for Generalization | Eric Jangevjang.com
- [1802.05296] Stronger generalization bounds for deep nets via a compression approacharxiv.org