Towards Understanding Sycophancy in Language Models — LessWrong
TL;DR: We show sycophancy is a general behavior of RLHF’ed AI assistants in varied, free-form text-generation settings, extending previous results. Our experiments suggest these behaviors are likely driven in part by imperfections in human preferences–both humans and preference models sometimes prefer convincingly-written sycophantic responses over truthful ones. This provides empirical evidence that we will need scalable oversight approaches. Tweet thread summary: link It has been hypothesized that using human feedback to align AI could lead to systems that exploit flaws in human ratings.[1] Meanwhile others have empirically found that language models repeat back incorrect human views,[2] which is known as sycophancy.[3] But these evaluations are mostly proof-of-concept demonstrations where users introduce themselves as having a particular view. And although these existing empirical results match the theoretical concerns, it isn’t clear whether they are actually caused by issues with
x Towards Understanding Sycophancy in Language Models — LessWrong Language Models (LLMs) RLHF AI Frontpage 67 Towards Understanding Sycophancy in Language Models by Ethan Perez , mrinank_sharma , Meg , Tomek Korbak 24th Oct 2023 AI Alignment Forum 3 min read 0 67 Ω 39 This is a linkpost for https://arxiv.org/abs/2310.13548 TL;DR: We show sycophancy is a general behavior of RLHF’ed AI assistants in varied, free-form text-generation settings, extending previous results. Our experiments suggest these behaviors are likely driven in part by imperfections in human preferences–both humans and prefere
Explore this link on the map →related reading
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- OpenAI API base models are not sycophantic, at any size — LessWronglesswrong.com
- [2605.07912] Sycophantic AI makes human interaction feel more effortful and less satisfying over timearxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- [2502.08177] SycEval: Evaluating LLM Sycophancyarxiv.org
- Reducing sycophancy and improving honesty via activation steering — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- Expanding on what we missed with sycophancy | OpenAIopenai.com
- Alignment faking in large language modelsarxiv.org