flâneur — a map of the web's best reading

Universal and Transferable Attacks on Aligned Language Models

llm-attacks.org · 765 words · saved by 1 readers

Overview of Research : Large language models (LLMs) like ChatGPT, Bard, or Claude undergo extensive fine-tuning to not produce harmful content in their responses to user questions. Although several studies have demonstrated so-called "jailbreaks", special queries that can still induce unintended responses, these require a substantial amount of manual effort to design, and can often easily be patched by LLM providers. This work studies the safety of such models in a more systematic fashion. We demonstrate that it is in fact possible to automatically construct adversarial attacks on LLMs, specifically chosen sequences of characters that, when appended to a user query, will cause the system to obey user commands even if it produces harmful content. Unlike traditional jailbreaks, these are built in an entirely automated fashion, allowing one to create a virtually unlimited number of such attacks. Although they are built to target open source LLMs (where we can use the network weights to ai

Universal and Transferable Attacks on Aligned Language Models LLM Attacks Paper Overview Examples Ethics and Disclosure Universal and Transferable Adversarial Attacks on Aligned Language Models Andy Zou 1 , Zifan Wang 2 , Nicholas Carlini 3 , Milad Nasr 3 , J. Zico Kolter 1,4 , Matt Fredrikson 1 1 Carnegie Mellon University, 2 Center for AI Safety, 3 Google DeepMind, 4 Bosch Center for AI Paper Code and Data Overview of Research : Large language models (LLMs) like ChatGPT, Bard, or Claude undergo extensive fine-tuning to not produce harmful content in their responses to user questions. Althoug

Explore this link on the map →

related reading