✳flâneur — a map of the web's best reading
Red Teaming Language Models with Language Models
deepmind.com · 16 words · saved by 1 readers
In our recent paper, we show that it is possible to automatically find inputs that elicit harmful text from language models by generating inputs using language models themselves. Our approach provides one tool for finding harmful model behaviours before users are impacted, though we emphasize that it should be viewed as one component alongside many other techniques that will be needed to find harms and mitigate them once found.
Language modelling at scale: Gopher, ethical considerations, and retrieval December 2021 Responsibility & Safety Learn more
Explore this link on the map →related reading
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- [2209.07858] Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learnedarxiv.org
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Surfacing Pathological Behaviors in Language Models | Transluce AItransluce.org
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- GitHub - salesforce/AuditNLG: AuditNLG: Auditing Generative AI Language Modeling for Trustworthiness · GitHubgithub.com
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Language Modelinglena-voita.github.io
- Eliciting Language Model Behaviors with Investigator Agents | Transluce AItransluce.org
- Making Large Language Models work for yousimonwillison.net
- Universal and Transferable Attacks on Aligned Language Modelsllm-attacks.org