✳flâneur — a map of the web's best reading
Steering Llama-2 with contrastive activation additions — LessWrong
lesswrong.com · 6,617 words · saved by 1 readers
The effects of subtracting or adding a "sycophancy vector" to one bias term. TL;DR: By just adding e.g. a "sycophancy vector" to one bias term, we o…
x Steering Llama-2 with contrastive activation additions — LessWrong Corrigibility Activation Engineering MATS Program Power Seeking (AI) Sycophancy AI Frontpage 125 Steering Llama-2 with contrastive activation additions by Nina Panickssery , Wuschel Schulz , NickGabs , Meg , evhub , TurnTrout 2nd Jan 2024 AI Alignment Forum Linkpost for arxiv.org 9 min read 29 125 Ω 49 The effects of subtracting or adding a "sycophancy vector" to one bias term. TL;DR: By just adding e.g. a "sycophancy vector" to one bias term, we outperform supervised finetuning and few-shot prompting at steering completions
Explore this link on the map →related reading
- Steering GPT-2-XL by adding an activation vector — AI Alignment Forumalignmentforum.org
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- Reducing sycophancy and improving honesty via activation steering — LessWronglesswrong.com
- Steering GPT-2-XL by adding an activation vector — LessWronglesswrong.com
- Steering Might Stop Working Soon — LessWronglesswrong.com
- [2312.06681] Steering Llama 2 via Contrastive Activation Additionarxiv.org
- Implementing activation steering — LessWronglesswrong.com
- Dreaming Vectors: Gradient-descented steering vectors from Activation Oracles and using them to Red-Team AOs — LessWronglesswrong.com
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences — LessWronglesswrong.com