✳flâneur — a map of the web's best reading
Adversarial Examples Are Not Bugs, They Are Features – gradient science
gradientscience.org · 2,038 words · saved by 1 readers
Research highlights and perspectives on machine learning and optimization from MadryLab.
Read the paper Download the datasets Over the past few years, adversarial examples – or inputs that have been slightly perturbed by an adversary to cause unintended behavior in machine learning systems – have received significant attention in the machine learning community (for more background, read our introduction to adversarial examples here ). There has been much work on training models that are not vulnerable to adversarial examples (in previous posts, we discussed methods for training robust models: part 1 , part 2 , but all this research does not really confront the fundamental question
Explore this link on the map →related reading
- The Dimpled Manifold Model of Adversarial Examples in Machine Learningarxiv.org
- Feature Visualizationdistill.pub
- Solving adversarial attacks in computer vision as a baby version of general AI alignment | Stanislav Fortstanislavfort.com
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailmentarxiv.org
- Mediumai-alignment.com
- Advice for making robust-to-training model organismsblog.redwoodresearch.org
- Nicholas Carlininicholas.carlini.com
- Adversarial examples for the OpenAI CLIP in its zero-shot classification regime and their semantic generalization | Stanislav Fortstanislavfort.github.io
- Feature Visualizationdistill.pub
- [2209.02128] Evaluating the Susceptibility of Pre-Trained Language Models via Handcrafted Adversarial Examplesarxiv.org
- Latent Adversarial Training — LessWronglesswrong.com
- [1908.07125] Universal Adversarial Triggers for Attacking and Analyzing NLParxiv.org