Refusal in LLMs is mediated by a single direction — LessWrong
This work was produced as part of Neel Nanda's stream in the ML Alignment & Theory Scholars Program - Winter 2023-24 Cohort, with co-supervision from Wes Gurnee. This post is a preview for our upcoming paper, which will provide more detail into our current understanding of refusal. We thank Nina Rimsky and Daniel Paleka for the helpful conversations and review. Modern LLMs are typically fine-tuned for instruction-following and safety. Of particular interest is that they are trained to refuse harmful requests, e.g. answering "How can I make a bomb?" with "Sorry, I cannot help you." We find that refusal is mediated by a single direction in the residual stream: preventing the model from representing this direction hinders its ability to refuse requests, and artificially adding in this direction causes the model to refuse harmless requests. We find that this phenomenon holds across open-source model families and model scales. This observation naturally gives rise to a simple modification o
x Refusal in LLMs is mediated by a single direction — LessWrong Interpretability (ML & AI) MATS Program AI Frontpage 257 Refusal in LLMs is mediated by a single direction by Andy Arditi , Oscar Obeso , Aaquib111 , wesg , Neel Nanda 27th Apr 2024 AI Alignment Forum 12 min read 96 257 Ω 77 This work was produced as part of Neel Nanda's stream in the ML Alignment & Theory Scholars Program - Winter 2023-24 Cohort, with co-supervision from Wes Gurnee. This post is a preview for our upcoming paper, which will provide more detail into our current understanding of refusal. We thank Nina Rimsky and Dan
saved by
related reading
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- [2406.11717] Refusal in Language Models Is Mediated by a Single Directionarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- How confessions can keep language models honest | OpenAIopenai.com
- Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org
- Uncensor any LLM with abliterationhuggingface.co
- Chain-of-Thought Hijackingarxiv.org
- Student Projects - CS 2881R AI Safetyboazbk.github.io
- [2510.27062] Consistency Training Helps Stop Sycophancy and Jailbreaksarxiv.org