flâneur — a map of the web's best reading

Towards Understanding Sycophancy in Language Models — LessWrong

lesswrong.com · 989 words · saved by 1 readers

TL;DR: We show sycophancy is a general behavior of RLHF’ed AI assistants in varied, free-form text-generation settings, extending previous results. Our experiments suggest these behaviors are likely driven in part by imperfections in human preferences–both humans and preference models sometimes prefer convincingly-written sycophantic responses over truthful ones. This provides empirical evidence that we will need scalable oversight approaches. Tweet thread summary: link It has been hypothesized that using human feedback to align AI could lead to systems that exploit flaws in human ratings.[1] Meanwhile others have empirically found that language models repeat back incorrect human views,[2] which is known as sycophancy.[3] But these evaluations are mostly proof-of-concept demonstrations where users introduce themselves as having a particular view. And although these existing empirical results match the theoretical concerns, it isn’t clear whether they are actually caused by issues with

x Towards Understanding Sycophancy in Language Models — LessWrong Language Models (LLMs) RLHF AI Frontpage 67 Towards Understanding Sycophancy in Language Models by Ethan Perez , mrinank_sharma , Meg , Tomek Korbak 24th Oct 2023 AI Alignment Forum 3 min read 0 67 Ω 39 This is a linkpost for https://arxiv.org/abs/2310.13548 TL;DR: We show sycophancy is a general behavior of RLHF’ed AI assistants in varied, free-form text-generation settings, extending previous results. Our experiments suggest these behaviors are likely driven in part by imperfections in human preferences–both humans and prefere

Explore this link on the map →

related reading