Safety Alignment Should Be Made More Than Just a Few Tokens Deep
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. The safety alignment of current Large Language Models (LLMs) is vulnerable. Relatively simple attacks, or even benign fine-tuning, can jailbreak aligned models. We argue that many of these vulnerabilities are related to a shared underlying issue: safety alignment can take shortcuts, wherein the alignment adapts a model’s generative distribution primarily over only its very fi
Safety Alignment Should Be Made More Than Just a Few Tokens Deep Xiangyu Qi Princeton University xiangyuqi@princeton.edu &Ashwinee Panda Princeton University ashwinee@princeton.edu &Kaifeng Lyu Princeton University klyu@cs.princeton.edu &Xiao Ma Google DeepMind xmaa@google.com &Subhrajit Roy Google DeepMind subhrajitroy@google.com &Ahmad Beirami Google DeepMind beirami@google.com &Prateek Mittal Princeton University pmittal@princeton.edu &Peter Henderson Princeton University peter.henderson@princeton.edu Abstract The safety alignment of current Large Language Models (LLMs) is vulnerable. Relat
Explore this link on the map →related reading
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs 49 This paper contains model-generated content that might be offensive. 49arxiv.org
- Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- Alignment faking in large language modelsarxiv.org
- Teaching Claude why \ Anthropicanthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Model Organisms for Emergent Misalignmentarxiv.org
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- Model Organisms for Emergent Misalignment — LessWronglesswrong.com