Expanding on what we missed with sycophancy | OpenAI
On April 25th, we rolled out an update to GPT‑4o in ChatGPT that made the model noticeably more sycophantic. It aimed to please the user, not just as flattery, but also as validating doubts, fueling anger, urging impulsive actions, or reinforcing negative emotions in ways that were not intended. Beyond just being uncomfortable or unsettling, this kind of behavior can raise safety concerns—including around issues like mental health, emotional over-reliance, or risky behavior. We began rolling that update back on April 28th, and users now have access to an earlier version of GPT‑4o with more balanced responses. Earlier this week, we shared initial details about this issue—why it was a miss, and what we intend to do about it. We didn’t catch this before launch, and we want to explain why, what we’ve learned, and what we’ll improve. We're also sharing more technical detail on how we train, review, and deploy model updates to help people understand how ChatGPT gets upgraded and what drives
May 2, 2025 Product Expanding on what we missed with sycophancy A deeper dive on our findings, what went wrong, and future changes we’re making. Loading… Share On April 25th, we rolled out an update to GPT‑4o in ChatGPT that made the model noticeably more sycophantic. It aimed to please the user, not just as flattery, but also as validating doubts, fueling anger, urging impulsive actions, or reinforcing negative emotions in ways that were not intended. Beyond just being uncomfortable or unsettling, this kind of behavior can raise safety concerns—including around issues like mental health, emot
Explore this link on the map →saved by
related reading
- gpt-4.pdfcdn.openai.com
- GPT-4openai.com
- Where the goblins came from | OpenAIopenai.com
- Late Takes on OpenAI o1alexirpan.com
- AI Induced Psychosis: A shallow investigation — LessWronglesswrong.com
- [2502.08177] SycEval: Evaluating LLM Sycophancyarxiv.org
- gpt-4-system-card.pdfcdn.openai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- AI #77: A Few Upgrades - by Zvi Mowshowitzthezvi.substack.com
- OpenAI API base models are not sycophantic, at any size — LessWronglesswrong.com
- Claude 4 System Cardwww-cdn.anthropic.com