The Dimpled Manifold Model of Adversarial Examples in Machine Learning
The extreme fragility of deep neural networks when presented with tiny perturbations in their inputs was independently discovered by several research groups in 2013, but in spite of enormous effort these adversarial examples remained a baffling phenomenon with no clear explanation. In this paper we introduce a new conceptual framework (which we call the Dimpled Manifold Model) which provides a simple explanation for why adversarial examples exist, why their perturbations have such tiny norms, why these perturbations look like random noise, and why a network which was adversarially trained with incorrectly labeled images can still correctly classify test images. In the last part of the paper we describe the results of numerous experiments which strongly support this new model, and in particular our assertion that adversarial perturbations are roughly perpendicular to the low dimensional manifold which contains all the training examples.
The extreme fragility of deep neural networks when presented with tiny perturbations in their inputs was independently discovered by several research groups in 2013, but in spite of enormous effort these adversarial examples remained a baffling phenomenon with no clear explanation. In this paper we introduce a new conceptual framework (which we call the Dimpled Manifold Model) which provides a simple explanation for why adversarial examples exist, why their perturbations have such tiny norms, why these perturbations look like random noise, and why a network which was adversarially trained with
Explore this link on the map →related reading
- Adversarial Examples Are Not Bugs, They Are Features – gradient sciencegradientscience.org
- Neural Networks, Manifolds, and Topology -- colah's blogcolah.github.io
- A Recipe for Training Neural Networkskarpathy.github.io
- Feature Visualizationdistill.pub
- A Recipe for Training Neural Networkskarpathy.github.io
- The Little Book of Deep Learningfleuret.org
- Solving adversarial attacks in computer vision as a baby version of general AI alignment | Stanislav Fortstanislavfort.com
- Nicholas Carlininicholas.carlini.com
- [2503.02113] Deep Learning is Not So Mysterious or Differentarxiv.org
- Feature Visualizationdistill.pub
- arxiv.org/pdf/1805.08522arxiv.org
- Theoretical Motivations for Deep Learning | Rinu Boneyrinuboney.github.io