Reducing sycophancy and improving honesty via activation steering — LessWrong
Produced as part of the SERI ML Alignment Theory Scholars Program - Summer 2023 Cohort, under the mentorship of Evan Hubinger. I generate an activation steering vector using Anthropic's sycophancy dataset and then find that this can be used to increase or reduce performance on TruthfulQA, indicating a common direction between sycophancy on questions of opinion and untruthfulness on questions relating to common misconceptions. I think this could be a promising research direction to understand dishonesty in language models better. Sycophancy in LLMs refers to the behavior when a model tells you what it thinks you want to hear / would approve of instead of what it internally represents as the truth. Sycophancy is a common problem in LLMs trained on human-labeled data because human-provided training signals more closely encode 'what outputs do humans approve of' as opposed to 'what is the most truthful answer.' According to Anthropic's paper Discovering Language Model Behaviors with Model
x Reducing sycophancy and improving honesty via activation steering — LessWrong Activation Engineering Language Models (LLMs) MATS Program Sycophancy AI Frontpage 122 Reducing sycophancy and improving honesty via activation steering by Nina Panickssery 28th Jul 2023 AI Alignment Forum 11 min read 18 122 Ω 52 Produced as part of the SERI ML Alignment Theory Scholars Program - Summer 2023 Cohort, under the mentorship of Evan Hubinger. I generate an activation steering vector using Anthropic's sycophancy dataset and then find that this can be used to increase or reduce performance on TruthfulQA,
Explore this link on the map →related reading
- Steering Llama-2 with contrastive activation additions — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- Activation Steering in 2026: A Practitioner's Field Guide | Subhadip Mitrasubhadipmitra.com
- Towards Understanding Sycophancy in Language Models — LessWronglesswrong.com
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- Expanding on what we missed with sycophancy | OpenAIopenai.com
- Steering GPT-2-XL by adding an activation vector — AI Alignment Forumalignmentforum.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Steering Might Stop Working Soon — LessWronglesswrong.com
- confessions_paper.pdfcdn.openai.com
- [2605.07912] Sycophantic AI makes human interaction feel more effortful and less satisfying over timearxiv.org
- OpenAI API base models are not sycophantic, at any size — LessWronglesswrong.com