Universal and Transferable Attacks on Aligned Language Models
Overview of Research : Large language models (LLMs) like ChatGPT, Bard, or Claude undergo extensive fine-tuning to not produce harmful content in their responses to user questions. Although several studies have demonstrated so-called "jailbreaks", special queries that can still induce unintended responses, these require a substantial amount of manual effort to design, and can often easily be patched by LLM providers. This work studies the safety of such models in a more systematic fashion. We demonstrate that it is in fact possible to automatically construct adversarial attacks on LLMs, specifically chosen sequences of characters that, when appended to a user query, will cause the system to obey user commands even if it produces harmful content. Unlike traditional jailbreaks, these are built in an entirely automated fashion, allowing one to create a virtually unlimited number of such attacks. Although they are built to target open source LLMs (where we can use the network weights to ai
Universal and Transferable Attacks on Aligned Language Models LLM Attacks Paper Overview Examples Ethics and Disclosure Universal and Transferable Adversarial Attacks on Aligned Language Models Andy Zou 1 , Zifan Wang 2 , Nicholas Carlini 3 , Milad Nasr 3 , J. Zico Kolter 1,4 , Matt Fredrikson 1 1 Carnegie Mellon University, 2 Center for AI Safety, 3 Google DeepMind, 4 Bosch Center for AI Paper Code and Data Overview of Research : Large language models (LLMs) like ChatGPT, Bard, or Claude undergo extensive fine-tuning to not produce harmful content in their responses to user questions. Althoug
Explore this link on the map →related reading
- [2307.15043] Universal and Transferable Adversarial Attacks on Aligned Language Modelsarxiv.org
- Adversarial Attacks on Aligned Language Models | Gray Swan Researchgrayswan.ai
- Adversarial Attacks on LLMs | Lil'Loglilianweng.github.io
- Nicholas Carlininicholas.carlini.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- [2602.04899] Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoningarxiv.org
- The lethal trifecta for AI agents: private data, untrusted content, and external communicationsimonwillison.net
- Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- Catching up on the weird world of LLMssimonwillison.net