flâneur — a map of the web's best reading

Red Teaming Language Models with Language Models

deepmind.com · 16 words · saved by 1 readers

In our recent paper, we show that it is possible to automatically find inputs that elicit harmful text from language models by generating inputs using language models themselves. Our approach provides one tool for finding harmful model behaviours before users are impacted, though we emphasize that it should be viewed as one component alongside many other techniques that will be needed to find harms and mitigate them once found.

Language modelling at scale: Gopher, ethical considerations, and retrieval December 2021 Responsibility & Safety Learn more

Explore this link on the map →

related reading