Adversarial Attacks on Aligned Language Models | Gray Swan Research
Deep learning models, though they perform well in many circumstances, are known to be vulnerable to a class of manipulation known as "adversarial attacks." These attacks manipulate inputs to the model so as to change its intended behavior; a common illustration of such attacks manipulates the pixels of a image slightly so that an image of one object (say, a pig), can be classified as another unrelated object, even though the images look indistinguishable to the human eye.
Adversarial Attacks on Aligned Language Models | Gray Swan Research Join the Jump into the Arena Solutions Partners Research Resources Company Log In Schedule a Demo Solutions Partners Research Resources Company Research Monitoring & Evaluation Adversarial Attacks on Aligned Language Models In July 2023, we published the first-ever automated jailbreaking method on large language models (LLMs) and exposed their susceptibility to adversarial attacks. By demonstrating that specific character sequences could bypass sophisticated safeguards, we highlighted a significant vulnerability that has urgen
Explore this link on the map →related reading
- Adversarial Attacks on LLMs | Lil'Loglilianweng.github.io
- [2307.15043] Universal and Transferable Adversarial Attacks on Aligned Language Modelsarxiv.org
- Universal and Transferable Attacks on Aligned Language Modelsllm-attacks.org
- Nicholas Carlininicholas.carlini.com
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- [2209.02128] Evaluating the Susceptibility of Pre-Trained Language Models via Handcrafted Adversarial Examplesarxiv.org
- Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samplesarxiv.org
- [2304.11082] Fundamental Limitations of Alignment in Large Language Modelsarxiv.org