Steering Might Stop Working Soon — LessWrong
lesswrong.com · 3,250 words · saved by 2 readers
Steering LLMs with single-vector methods might break down soon, and by soon I mean soon enough that if you're working on steering, you should start p…
x Steering Might Stop Working Soon — LessWrong AI Frontpage 68 Steering Might Stop Working Soon by J Bostock 5th Apr 2026 5 min read 13 68 Steering LLMs with single-vector methods might break down soon, and by soon I mean soon enough that if you're working on steering, you should start planning for it failing now . This is particularly important for things like steering as a mitigation against eval-awareness. Steering Humans I have a strong intuition that we will not be able to steer a superintelligence very effectively, partially for the same reason that you probably can't steer a human very
saved by
related reading
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- [2605.03907] Steer Like the LLM: Activation Steering that Mimics Promptingarxiv.org
- Reproducing steering against evaluation awareness in a large open-weight model — LessWronglesswrong.com
- Steering GPT-2-XL by adding an activation vector — AI Alignment Forumalignmentforum.org
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- AI in 2025: gestalt — LessWronglesswrong.com
- [2312.06681] Steering Llama 2 via Contrastive Activation Additionarxiv.org
- Steering Llama-2 with contrastive activation additions — LessWronglesswrong.com
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- Implementing activation steering — LessWronglesswrong.com