An SDE Framework for Adversarial Training, with Convergence and Robustness Analysis
Adversarial training has gained great popularity as one of the most effective defenses for deep neural networks against adversarial perturbations on data points. Consequently, research interests have grown in understanding the convergence and robustness of adversarial training. This paper considers the min-max game of adversarial training by alternating stochastic gradient descent. It approximates the training process with a continuous-time stochastic-differential-equation (SDE). In particular, the error bound and convergence analysis is established. This SDE framework allows direct comparison between adversarial training and stochastic gradient descent; and confirms analytically the robustness of adversarial training from a (new) gradient-flow viewpoint. This analysis is then corroborated via numerical studies. To demonstrate the versatility of this SDE framework for algorithm design and parameter tuning, a stochastic control problem is formulated for learning rate adjustment, where the advantage of adaptive learning rate over fixed learning rate in terms of training loss is demonstrated through numerical experiments.
Adversarial training has gained great popularity as one of the most effective defenses for deep neural networks against adversarial perturbations on data points. Consequently, research interests have grown in understanding the convergence and robustness of adversarial training. This paper considers the min-max game of adversarial training by alternating stochastic gradient descent. It approximates the training process with a continuous-time stochastic-differential-equation (SDE). In particular, the error bound and convergence analysis is established. This SDE framework allows direct comparison
Explore this link on the map →related reading
- Stochastic gradient descent - Wikipediaen.m.wikipedia.org
- The Decade of Deep Learning | Leo Gaobmk.sh
- [2101.12176] On the Origin of Implicit Regularization in Stochastic Gradient Descentarxiv.org
- The Little Book of Deep Learningfleuret.org
- [1512.04202] Preconditioned Stochastic Gradient Descentarxiv.org
- [1511.06251] Stochastic modified equations and adaptive stochastic gradient algorithmsarxiv.org
- Neural SDE: Stabilizing Neural ODE Networks with Stochastic Noisearxiv.org
- AdaGrad - Cornell University Computational Optimization Open Textbook - Optimization Wikioptimization.cbe.cornell.edu
- [2605.01172] A Theory of Generalization in Deep Learningarxiv.org
- Notes on the Origin of Implicit Regularization in SGDinference.vc
- 1806.07366arxiv.org
- Yang Songyang-song.net