flâneur — a map of the web's best reading

Steering Llama-2 with contrastive activation additions — LessWrong

lesswrong.com · 6,617 words · saved by 1 readers

The effects of subtracting or adding a "sycophancy vector" to one bias term. TL;DR: By just adding e.g. a "sycophancy vector" to one bias term, we o…

x Steering Llama-2 with contrastive activation additions — LessWrong Corrigibility Activation Engineering MATS Program Power Seeking (AI) Sycophancy AI Frontpage 125 Steering Llama-2 with contrastive activation additions by Nina Panickssery , Wuschel Schulz , NickGabs , Meg , evhub , TurnTrout 2nd Jan 2024 AI Alignment Forum Linkpost for arxiv.org 9 min read 29 125 Ω 49 The effects of subtracting or adding a "sycophancy vector" to one bias term. TL;DR: By just adding e.g. a "sycophancy vector" to one bias term, we outperform supervised finetuning and few-shot prompting at steering completions

Explore this link on the map →

related reading