Adversarial Examples Are Not Bugs, They Are Features – gradient science
gradientscience.org · 2,038 words · saved by 1 readers
Research highlights and perspectives on machine learning and optimization from MadryLab.
Read the paper Download the datasets Over the past few years, adversarial examples – or inputs that have been slightly perturbed by an adversary to cause unintended behavior in machine learning systems – have received significant attention in the machine learning community (for more background, read our introduction to adversarial examples here ). There has been much work on training models that are not vulnerable to adversarial examples (in previous posts, we discussed methods for training robust models: part 1 , part 2 , but all this research does not really confront the fundamental question
related reading
- The Dimpled Manifold Model of Adversarial Examples in Machine Learningarxiv.org
- Some Lessons from Adversarial Machine Learning | FAR.AIfar.ai
- Feature Visualizationdistill.pub
- [1610.00768] Technical Report on the CleverHans v2.1.0 Adversarial Examples Libraryarxiv.org
- Solving adversarial attacks in computer vision as a baby version of general AI alignment | Stanislav Fortstanislavfort.com
- Your Out-of-Distribution Detection Method is Not Robust!arxiv.org
- Nicholas Carlininicholas.carlini.com
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailmentarxiv.org
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- Advice for making robust-to-training model organismsblog.redwoodresearch.org
- Mediumai-alignment.com
- Understanding deep learning requires rethinking generalizationarxiv.org