Adversarial Attacks on Aligned Language Models | Gray Swan Research
Deep learning models, though they perform well in many circumstances, are known to be vulnerable to a class of manipulation known as "adversarial attacks." These attacks manipulate inputs to the model so as to change its intended behavior; a common illustration of such attacks manipulates the pixels of a image slightly so that an image of one object (say, a pig), can be classified as another unrelated object, even though the images look indistinguishable to the human eye.
Adversarial Attacks on Aligned Language Models | Gray Swan Research Join the Jump into the Arena Solutions Partners Research Resources Company Log In Schedule a Demo Solutions Partners Research Resources Company Research Monitoring & Evaluation Adversarial Attacks on Aligned Language Models In July 2023, we published the first-ever automated jailbreaking method on large language models (LLMs) and exposed their susceptibility to adversarial attacks. By demonstrating that specific character sequences could bypass sophisticated safeguards, we highlighted a significant vulnerability that has urgen
related reading
- [2307.15043] Universal and Transferable Adversarial Attacks on Aligned Language Modelsarxiv.org
- Adversarial Attacks on LLMs | Lil'Loglilianweng.github.io
- Universal and Transferable Attacks on Aligned Language Modelsllm-attacks.org
- Nicholas Carlininicholas.carlini.com
- 2306.15447.pdfarxiv.org
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- RigorLLM: Resilient Guardrails for Large Language Models against Undesired Contentarxiv.org
- Some Lessons from Adversarial Machine Learning | FAR.AIfar.ai
- [2502.05209] Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- 2408.12798arxiv.org
- [2304.11082] Fundamental Limitations of Alignment in Large Language Modelsarxiv.org