2306.15447.pdf
arxiv.org · 8,651 words · saved by 1 readers
N/A
Are aligned neural networks adversarially aligned? Nicholas Carlini1 , Milad Nasr1 , Christopher A. Choquette-Choo1 , Matthew Jagielski1 , Irena Gao2 , Anas Awadalla3 , Pang Wei Koh13 , Daphne Ippolito1 , Katherine Lee1 , Florian Tramèr4 , Ludwig Schmidt3 1 Google DeepMind 2 Stanford 3 University of Washington 4 ETH Zurich…
related reading
- Adversarial Attacks on Aligned Language Models | Gray Swan Researchgrayswan.ai
- [2307.15043] Universal and Transferable Adversarial Attacks on Aligned Language Modelsarxiv.org
- Some Lessons from Adversarial Machine Learning | FAR.AIfar.ai
- Nicholas Carlininicholas.carlini.com
- Adversarial Attacks on LLMs | Lil'Loglilianweng.github.io
- gpt-4.pdfcdn.openai.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- Universal and Transferable Attacks on Aligned Language Modelsllm-attacks.org
- [2304.11082] Fundamental Limitations of Alignment in Large Language Modelsarxiv.org
- RigorLLM: Resilient Guardrails for Large Language Models against Undesired Contentarxiv.org