Steering Llama-2 with contrastive activation additions — LessWrong
lesswrong.com · 6,617 words · saved by 1 readers
The effects of subtracting or adding a "sycophancy vector" to one bias term. TL;DR: By just adding e.g. a "sycophancy vector" to one bias term, we o…
x Steering Llama-2 with contrastive activation additions — LessWrong Corrigibility Activation Engineering MATS Program Power Seeking (AI) Sycophancy AI Frontpage 125 Steering Llama-2 with contrastive activation additions by Nina Panickssery , Wuschel Schulz , NickGabs , Meg , evhub , TurnTrout 2nd Jan 2024 AI Alignment Forum Linkpost for arxiv.org 9 min read 29 125 Ω 49 The effects of subtracting or adding a "sycophancy vector" to one bias term. TL;DR: By just adding e.g. a "sycophancy vector" to one bias term, we outperform supervised finetuning and few-shot prompting at steering completions
related reading
- Steering GPT-2-XL by adding an activation vector — AI Alignment Forumalignmentforum.org
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- [2312.06681] Steering Llama 2 via Contrastive Activation Additionarxiv.org
- Reducing sycophancy and improving honesty via activation steering — LessWronglesswrong.com
- Steering GPT-2-XL by adding an activation vector — LessWronglesswrong.com
- Steering Might Stop Working Soon — LessWronglesswrong.com
- Implementing activation steering — LessWronglesswrong.com
- [2605.03907] Steer Like the LLM: Activation Steering that Mimics Promptingarxiv.org
- Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs — LessWronglesswrong.com
- [2607.14111] Introspection Fine-Tuning (IFT): Training Small LLMs to Introspectarxiv.org
- 2308.03958arxiv.org
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org