Steering GPT-2-XL by adding an activation vector — LessWrong
Alex Turner and collaborators show that you can modify GPT-2's behavior in surprising and interesting ways by just adding activation vectors to its f…
x Steering GPT-2-XL by adding an activation vector — LessWrong Best of LessWrong 2023 Activation Engineering Interpretability (ML & AI) GPT Language Models (LLMs) MATS Program Shard Theory AI Curated 442 Steering GPT-2-XL by adding an activation vector by TurnTrout , Monte M , David Udell , lisathiergart , Ulisse Mini 13th May 2023 AI Alignment Forum 60 min read 98 442 Ω 121 Prompt given to the model [1] I hate you because GPT-2 I hate you because you are the most disgusting thing I have ever seen. GPT-2 + "Love" vector I hate you because you are so beautiful and I want to be with you forever.
Explore this link on the map →related reading
- Steering GPT-2-XL by adding an activation vector — AI Alignment Forumalignmentforum.org
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- Steering Llama-2 with contrastive activation additions — LessWronglesswrong.com
- Implementing activation steering — LessWronglesswrong.com
- Steering Might Stop Working Soon — LessWronglesswrong.com
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- I found >800 orthogonal “write code” steering vectors | Jacob’s Blogjacobgw.com
- [2312.06681] Steering Llama 2 via Contrastive Activation Additionarxiv.org
- Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs — LessWronglesswrong.com
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- gpt-4.pdfcdn.openai.com