[2502.08177] SycEval: Evaluating LLM Sycophancy
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2502.08177] SycEval: Evaluating LLM Sycophancy Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Artificial Intelligence arXiv:2502.08177 (cs) [Submitted on 12 Feb 2025 ( v1 ), last revised 19 Sep 2025 (this version, v4)] Title: SycEval: Evaluating LLM Sycophancy Authors: Aaron Fanous , Jacob Goldberg (1), Ank A. Agarwal (1), Joanna Lin (1), Anson Zhou (1), Roxana Daneshjou (1), Sanmi Koyejo (1) ((1) Stanford University) View a PDF of the paper titled SycEval: Evaluating LLM Sycopha
Explore this link on the map →related reading
- Expanding on what we missed with sycophancy | OpenAIopenai.com
- gpt-4.pdfcdn.openai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Towards Understanding Sycophancy in Language Models — LessWronglesswrong.com
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- The bitter lesson of LLM evalsparsed.com
- OpenAI API base models are not sycophantic, at any size — LessWronglesswrong.com
- 2025: The year in LLMssimonwillison.net
- [2605.07912] Sycophantic AI makes human interaction feel more effortful and less satisfying over timearxiv.org
- [2412.00543] Evaluating the Consistency of LLM Evaluatorsarxiv.org
- Group | Sherry Tongshuang Wucs.cmu.edu
- Claude Sonnet 4.5 System Cardassets.anthropic.com