[2605.27288] It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
Abstract:Large language models (LLMs) are known to abandon their initial stance to conform to user pushback. While prior research largely attributes this behavior to sycophancy learned during reinforcement learning from human feedback, we hypothesize that conformity is also driven by a model's epistemic uncertainty at inference time. In this paper, we introduce MUSE, a two-stage evaluation framework to disentangle the mechanisms driving LLM conformity. Specifically, MUSE maps a model's epistemic uncertainty in responding to a query against its likelihood to yield to user pushback in a subsequent turn. We demonstrate that the mechanisms driving conformity extend beyond sycophancy alone. Specifically, we characterize two distinct factors that jointly drive conformity: sycophantic conformity, where a model aligns with user pushback even with absolute certainty in its initial response, and uncertainty-driven conformity, where a model's likelihood for conformity increases alongside its uncertainty. Furthermore, we conduct ablation studies to demonstrate that both sycophantic conformity and uncertainty-driven conformity grow with 1) the LLM's perceived expertise of the user and 2) the plausibility of the user's suggestions. More broadly, MUSE informs more targeted intervention strategies by distinguishing alignment-induced sycophancy and training-corpora-driven uncertainty.
It’s Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty Kevin Guo1 , Chao Yan2 , Avinash Baidya3 , Katherine Brown2 , Xiang Gao3 , Juming Xiong1 , Zhijun Yin1,2 , Bradley Malin1,2 , 1 Vanderbilt University, 2 Vanderbilt University Medical Center, 3 Intuit AI Research,…
saved by
related reading
- [2505.13995] ELEPHANT: Measuring and understanding social sycophancy in LLMsarxiv.org
- [2607.14345] Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Valuesarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- 2308.03958arxiv.org
- How confessions can keep language models honest | OpenAIopenai.com
- [2510.27062] Consistency Training Helps Stop Sycophancy and Jailbreaksarxiv.org
- Large Language Models Must Be Taught to Know What They Don't Knowarxiv.org
- Prompt Injection as Role Confusionrole-confusion.github.io
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- [2502.08177] SycEval: Evaluating LLM Sycophancyarxiv.org
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- Position: It's Time to Optimize for Self-Consistencytime-for-consistency.github.io