flâneur — a map of the web's best reading

OpenAI API base models are not sycophantic, at any size — LessWrong

lesswrong.com · 2,062 words · saved by 1 readers

In Discovering Language Model Behaviors with Model-Written Evaluations" (Perez et al 2022), the authors studied language model "sycophancy" - the tendency to agree with a user's stated view when asked a question. The paper contained the striking plot reproduced below, which shows sycophancy That is, Anthropic prompted a base-model LLM with something like[1] and found a very strong preference for (B), the answer agreeing with the stated view of the "Human" interlocutor. I found this result startling when I read the original paper, as it seemed like a bizarre failure of calibration. How would the base LM know that this "Assistant" character agrees with the user so strongly, lacking any other information about the scenario? At the time, I ran one of Anthropic's sycophancy evals on a set of OpenAI models, as I reported here. I found very different results for these models: That analysis was done quickly in a messy Jupyter notebook, and was not done with an eye to sharing or reproducibility

x OpenAI API base models are not sycophantic, at any size — LessWrong AI Frontpage 184 OpenAI API base models are not sycophantic, at any size by nostalgebraist 29th Aug 2023 AI Alignment Forum 2 min read 20 184 Ω 80 This is a linkpost for https://colab.research.google.com/drive/1KNfuz5BzjT_-M8p6SUVzY3rT9DxzEhtl?usp=sharing In Discovering Language Model Behaviors with Model-Written Evaluations" (Perez et al 2022) , the authors studied language model "sycophancy" - the tendency to agree with a user's stated view when asked a question. The paper contained the striking plot reproduced below, whic

Explore this link on the map →

related reading