Solving adversarial attacks in computer vision as a baby version of general AI alignment | Stanislav Fort
I spent the last few months trying to tackle the problem of adversarial attacks in computer vision from the ground up. The results of this effort are written up in our new paper Ensemble everything everywhere: Multi-scale aggregation for adversarial robustness (explainer on X/Twitter). Taking inspiration from biology, we reached state-of-the-art or above state-of-the-art robustness at 100x – 1000x less compute, got human-understandable interpretability for free, turned classifiers into generators, and designed transferable adversarial attacks on closed-source (v)LLMs such as GPT-4 or Claude 3. I strongly believe that there is a compelling case for devoting serious attention to solving the problem of adversarial robustness in computer vision, and I try to draw an analogy to the alignment of general AI systems here. In this post, I argue that the problem of adversarial attacks in computer vision is in many ways analogous to the larger task of general AI alignment. In both cases, we are t
Stanislav Fort ( Twitter , website , and Google Scholar ) I spent the last few months trying to tackle the problem of adversarial attacks in computer vision from the ground up. The results of this effort are written up in our new paper Ensemble everything everywhere: Multi-scale aggregation for adversarial robustness ( explainer on X/Twitter ). Taking inspiration from biology, we reached state-of-the-art or above state-of-the-art robustness at 100x – 1000x less compute, got human-understandable interpretability for free, turned classifiers into generators, and designed transferable adversarial
saved by
related reading
- Some Lessons from Adversarial Machine Learning | FAR.AIfar.ai
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- The Limits of Computer Vision, and of Our Own | Harvard Medicine Magazinemagazine.hms.harvard.edu
- Adversarial Examples Are Not Bugs, They Are Features – gradient sciencegradientscience.org
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Mediumai-alignment.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- [2003.01690] Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksarxiv.org