2308.03958
arxiv.org · 7,162 words · saved by 2 readers
N/A
February 16, 2024 S IMPLE SYNTHETIC DATA REDUCES SYCOPHANCY IN LARGE LANGUAGE MODELS Jerry Wei Da Huang Yifeng Lu Denny Zhou Quoc V. Le Google DeepMind A BSTRACT arXiv:2308.03958v2 [cs.CL] 15 Feb 2024 Sycophancy is an undesirable behavior where…
saved by
related reading
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2505.13995] ELEPHANT: Measuring and understanding social sycophancy in LLMsarxiv.org
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Towards Understanding Sycophancy in Language Models — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Expanding on what we missed with sycophancy | OpenAIopenai.com
- OpenAI API base models are not sycophantic, at any size — LessWronglesswrong.com
- Training language models to follow instructions with human feedback.pdfproceedings.neurips.cc
- [2502.08177] SycEval: Evaluating LLM Sycophancyarxiv.org
- [2303.17548] Whose Opinions Do Language Models Reflect?arxiv.org
- [2510.27062] Consistency Training Helps Stop Sycophancy and Jailbreaksarxiv.org