flâneur — a map of the web's best reading

Reducing sycophancy and improving honesty via activation steering — LessWrong

lesswrong.com · 4,893 words · saved by 1 readers

Produced as part of the SERI ML Alignment Theory Scholars Program - Summer 2023 Cohort, under the mentorship of Evan Hubinger. I generate an activation steering vector using Anthropic's sycophancy dataset and then find that this can be used to increase or reduce performance on TruthfulQA, indicating a common direction between sycophancy on questions of opinion and untruthfulness on questions relating to common misconceptions. I think this could be a promising research direction to understand dishonesty in language models better. Sycophancy in LLMs refers to the behavior when a model tells you what it thinks you want to hear / would approve of instead of what it internally represents as the truth. Sycophancy is a common problem in LLMs trained on human-labeled data because human-provided training signals more closely encode 'what outputs do humans approve of' as opposed to 'what is the most truthful answer.' According to Anthropic's paper Discovering Language Model Behaviors with Model

x Reducing sycophancy and improving honesty via activation steering — LessWrong Activation Engineering Language Models (LLMs) MATS Program Sycophancy AI Frontpage 122 Reducing sycophancy and improving honesty via activation steering by Nina Panickssery 28th Jul 2023 AI Alignment Forum 11 min read 18 122 Ω 52 Produced as part of the SERI ML Alignment Theory Scholars Program - Summer 2023 Cohort, under the mentorship of Evan Hubinger. I generate an activation steering vector using Anthropic's sycophancy dataset and then find that this can be used to increase or reduce performance on TruthfulQA,

Explore this link on the map →

related reading