[2605.03907] Steer Like the LLM: Activation Steering that Mimics Prompting
Abstract:Large language models can be steered at inference time through prompting or activation interventions, but activation steering methods often underperform compared to prompt-based approaches. We propose a framework that formulates prompt steering as a form of activation steering and investigates whether distilling successful prompt steering behavior into simpler, interpretable models can close this gap. Our analysis reveals that popular activation steering methods are not faithful to the mechanics of prompt steering, which applies strong interventions on some tokens while barely affecting others. Based on these insights, we introduce Prompt Steering Replacement (PSR) models that estimate token-specific steering coefficients from the activations themselves and are trained to imitate prompt-based interventions. Experiments on three steering benchmarks across multiple language models show that PSR models outperform existing activation steering methods, especially when controlling for high-coherence completions, and also compare favorably to prompting on AxBench and persona steering.
Steer Like the LLM: Activation Steering that Mimics Prompting Geert Heyman 1 Frederik Vandeputte 1 Abstract ble to prompt injection attacks that override the intended Large language models can be steered at inference behavior (Anwar et al., 2024) and constructing prompts that time through prompting or…
saved by
related reading
- Steering Might Stop Working Soon — LessWronglesswrong.com
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- Steering GPT-2-XL by adding an activation vector — AI Alignment Forumalignmentforum.org
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- [2312.06681] Steering Llama 2 via Contrastive Activation Additionarxiv.org
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWronglesswrong.com
- [2607.14111] Introspection Fine-Tuning (IFT): Training Small LLMs to Introspectarxiv.org
- Implementing activation steering — LessWronglesswrong.com
- 2402.07927arxiv.org
- Steering Llama-2 with contrastive activation additions — LessWronglesswrong.com
- Productizing Large Language Modelsblog.replit.com
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org