Implementing activation steering — LessWrong
Produced as part of the SERI ML Alignment Theory Scholars Program - Autumn 2023 Cohort and while being an affiliate at PIBBSS in 2024. A thank you to @Jayjay and @fela for helpful comments on this draft. This blog post is an overview of different ways to implement activation steering with some of my takes on their pros and cons. See also this GitHub repository for my minimal implementations of the different approaches. The blog post is aimed at people who are new to activation/representation steering/engineering/editing. The idea is simple: we just add some vector to the internal model activations and thus influence the model output in a similar (but sometimes more effective way) to prompting. Example[1]: Imagine that some vector in the internal representations in some transformer layer encodes a direction associated with "Love". When you add this vector to the activations of some encoded sentence "I hate the world", you change the internal representation (and thus the meaning) to some
x Implementing activation steering — LessWrong Activation Engineering Interpretability (ML & AI) Language Models (LLMs) GPT MATS Program AI Frontpage 76 Implementing activation steering by Annah 5th Feb 2024 9 min read 8 76 Produced as part of the SERI ML Alignment Theory Scholars Program - Autumn 2023 Cohort and while being an affiliate at PIBBSS in 2024. A thank you to @Jayjay and @fela for helpful comments on this draft. This blog post is an overview of different ways to implement activation steering with some of my takes on their pros and cons. See also this GitHub repository for my minima
Explore this link on the map →saved by
related reading
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- Steering GPT-2-XL by adding an activation vector — AI Alignment Forumalignmentforum.org
- Steering Llama-2 with contrastive activation additions — LessWronglesswrong.com
- Steering GPT-2-XL by adding an activation vector — LessWronglesswrong.com
- Steering Might Stop Working Soon — LessWronglesswrong.com
- Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs — LessWronglesswrong.com
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- I found >800 orthogonal “write code” steering vectors | Jacob’s Blogjacobgw.com
- [2312.06681] Steering Llama 2 via Contrastive Activation Additionarxiv.org