Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitra
I’ve been working with steering vectors for months. Here’s what actually works in practice, what fails in ways nobody warned me about, and the honest playbook for getting started.
TL;DR: Steering vectors are the most underrated tool in the LLM practitioner’s toolkit -and also the most oversold. They genuinely work for some behaviors (refusal, sentiment, formality). They genuinely fail for others (factual recall, complex reasoning). This is the guide I wish I’d had six months ago: what to steer, where to inject, how strong, and when to give up and try something else. The Promise and the Reality The pitch for activation steering is seductive. You take a pair of contrasting prompts (“be helpful” vs “refuse everything”), run them through the model, compute the mean activati
Explore this link on the map →saved by
related reading
- Steering Might Stop Working Soon — LessWronglesswrong.com
- Steering GPT-2-XL by adding an activation vector — AI Alignment Forumalignmentforum.org
- Implementing activation steering — LessWronglesswrong.com
- Steering Llama-2 with contrastive activation additions — LessWronglesswrong.com
- I found >800 orthogonal “write code” steering vectors | Jacob’s Blogjacobgw.com
- Steering GPT-2-XL by adding an activation vector — LessWronglesswrong.com
- Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs — LessWronglesswrong.com
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- Reproducing steering against evaluation awareness in a large open-weight model — LessWronglesswrong.com
- [2312.06681] Steering Llama 2 via Contrastive Activation Additionarxiv.org
- TurnTrout's shortform feed — LessWronglesswrong.com