Red Teaming Language Models with Language Models
In our recent paper, we show that it is possible to automatically find inputs that elicit harmful text from language models by generating inputs using language models themselves. Our approach provides one tool for finding harmful model behaviours before users are impacted, though we emphasize that it should be viewed as one component alongside many other techniques that will be needed to find harms and mitigate them once found.
February 7, 2022 Research Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, Geoffrey Irving In our recent paper, we show that it is possible to automatically find inputs that elicit harmful text from language models by generating inputs using language models themselves. Our approach provides one tool for finding harmful model behaviours before users are impacted, though we emphasize that it should be viewed as one component alongside many other techniques that will be needed to find harms and mitigate them once found. Large…
related reading
- [2209.07858] Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learnedarxiv.org
- gpt-4.pdfcdn.openai.com
- gpt-4-system-card.pdfcdn.openai.com
- The Dark Forest and Generative AImaggieappleton.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Nicholas Carlininicholas.carlini.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- Surfacing Pathological Behaviors in Language Models | Transluce AItransluce.org
- [2606.04071] Covert Influence Between Language Modelsarxiv.org
- Adversarial Attacks on LLMs | Lil'Loglilianweng.github.io