Universal and Transferable Attacks on Aligned Language Models
Overview of Research : Large language models (LLMs) like ChatGPT, Bard, or Claude undergo extensive fine-tuning to not produce harmful content in their responses to user questions. Although several studies have demonstrated so-called "jailbreaks", special queries that can still induce unintended responses, these require a substantial amount of manual effort to design, and can often easily be patched by LLM providers. This work studies the safety of such models in a more systematic fashion. We demonstrate that it is in fact possible to automatically construct adversarial attacks on LLMs, specifically chosen sequences of characters that, when appended to a user query, will cause the system to obey user commands even if it produces harmful content. Unlike traditional jailbreaks, these are built in an entirely automated fashion, allowing one to create a virtually unlimited number of such attacks. Although they are built to target open source LLMs (where we can use the network weights to ai
Universal and Transferable Attacks on Aligned Language Models LLM Attacks Paper Overview Examples Ethics and Disclosure Universal and Transferable Adversarial Attacks on Aligned Language Models Andy Zou 1 , Zifan Wang 2 , Nicholas Carlini 3 , Milad Nasr 3 , J. Zico Kolter 1,4 , Matt Fredrikson 1 1 Carnegie Mellon University, 2 Center for AI Safety, 3 Google DeepMind, 4 Bosch Center for AI Paper Code and Data Overview of Research : Large language models (LLMs) like ChatGPT, Bard, or Claude undergo extensive fine-tuning to not produce harmful content in their responses to user questions. Althoug
related reading
- [2307.15043] Universal and Transferable Adversarial Attacks on Aligned Language Modelsarxiv.org
- Adversarial Attacks on Aligned Language Models | Gray Swan Researchgrayswan.ai
- Adversarial Attacks on LLMs | Lil'Loglilianweng.github.io
- Nicholas Carlininicholas.carlini.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- RigorLLM: Resilient Guardrails for Large Language Models against Undesired Contentarxiv.org
- [2502.05209] Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- 2408.12798arxiv.org
- 2306.15447.pdfarxiv.org
- Simulated Users & Sad LLMs1a3orn.com
- Voices Across Registers: Corpus-Conditioned Vernacular Jailbreaks against Aligned LLMs via Fanfiction Subgenresarxiv.org