Refusal in LLMs is mediated by a single direction — LessWrong
This work was produced as part of Neel Nanda's stream in the ML Alignment & Theory Scholars Program - Winter 2023-24 Cohort, with co-supervision from Wes Gurnee. This post is a preview for our upcoming paper, which will provide more detail into our current understanding of refusal. We thank Nina Rimsky and Daniel Paleka for the helpful conversations and review. Modern LLMs are typically fine-tuned for instruction-following and safety. Of particular interest is that they are trained to refuse harmful requests, e.g. answering "How can I make a bomb?" with "Sorry, I cannot help you." We find that refusal is mediated by a single direction in the residual stream: preventing the model from representing this direction hinders its ability to refuse requests, and artificially adding in this direction causes the model to refuse harmless requests. We find that this phenomenon holds across open-source model families and model scales. This observation naturally gives rise to a simple modification o
x Refusal in LLMs is mediated by a single direction — LessWrong Interpretability (ML & AI) MATS Program AI Frontpage 257 Refusal in LLMs is mediated by a single direction by Andy Arditi , Oscar Obeso , Aaquib111 , wesg , Neel Nanda 27th Apr 2024 AI Alignment Forum 12 min read 96 257 Ω 77 This work was produced as part of Neel Nanda's stream in the ML Alignment & Theory Scholars Program - Winter 2023-24 Cohort, with co-supervision from Wes Gurnee. This post is a preview for our upcoming paper, which will provide more detail into our current understanding of refusal. We thank Nina Rimsky and Dan
Explore this link on the map →saved by
related reading
- [2406.11717] Refusal in Language Models Is Mediated by a Single Directionarxiv.org
- How confessions can keep language models honest | OpenAIopenai.com
- Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- Steering Might Stop Working Soon — LessWronglesswrong.com
- Alignment faking in large language modelsarxiv.org
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- Circuits Updates - April 2025transformer-circuits.pub