OpenAI API base models are not sycophantic, at any size — LessWrong
In Discovering Language Model Behaviors with Model-Written Evaluations" (Perez et al 2022), the authors studied language model "sycophancy" - the tendency to agree with a user's stated view when asked a question. The paper contained the striking plot reproduced below, which shows sycophancy That is, Anthropic prompted a base-model LLM with something like[1] and found a very strong preference for (B), the answer agreeing with the stated view of the "Human" interlocutor. I found this result startling when I read the original paper, as it seemed like a bizarre failure of calibration. How would the base LM know that this "Assistant" character agrees with the user so strongly, lacking any other information about the scenario? At the time, I ran one of Anthropic's sycophancy evals on a set of OpenAI models, as I reported here. I found very different results for these models: That analysis was done quickly in a messy Jupyter notebook, and was not done with an eye to sharing or reproducibility
x OpenAI API base models are not sycophantic, at any size — LessWrong AI Frontpage 184 OpenAI API base models are not sycophantic, at any size by nostalgebraist 29th Aug 2023 AI Alignment Forum 2 min read 20 184 Ω 80 This is a linkpost for https://colab.research.google.com/drive/1KNfuz5BzjT_-M8p6SUVzY3rT9DxzEhtl?usp=sharing In Discovering Language Model Behaviors with Model-Written Evaluations" (Perez et al 2022) , the authors studied language model "sycophancy" - the tendency to agree with a user's stated view when asked a question. The paper contained the striking plot reproduced below, whic
Explore this link on the map →related reading
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Towards Understanding Sycophancy in Language Models — LessWronglesswrong.com
- Expanding on what we missed with sycophancy | OpenAIopenai.com
- [2502.08177] SycEval: Evaluating LLM Sycophancyarxiv.org
- How confessions can keep language models honest | OpenAIopenai.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Reducing sycophancy and improving honesty via activation steering — LessWronglesswrong.com
- [2605.07912] Sycophantic AI makes human interaction feel more effortful and less satisfying over timearxiv.org
- Alignment faking in large language modelsarxiv.org
- [2303.17548] Whose Opinions Do Language Models Reflect?arxiv.org
- PostTrainBenchposttrainbench.com