Mysteries of mode collapse — LessWrong
Thanks to Ian McKenzie and Nicholas Dupuis, collaborators on a related project, for contributing to the ideas and experiments discussed in this post. Ian performed some of the random number experiments. Also thanks to Connor Leahy for feedback on a draft, and thanks to Evan Hubinger, Connor Leahy, Beren Millidge, Ethan Perez, Tomek Korbak, Garrett Baker, Leo Gao and various others at Conjecture, Anthropic, and OpenAI for useful discussions. This work was carried out while at Conjecture. I have received evidence from multiple credible sources that text-davinci-002 was not trained with RLHF. The rest of this post has not been corrected to reflect this update. Not much besides the title (formerly "Mysteries of mode collapse due to RLHF") is affected: just mentally substitute "mystery method" every time "RLHF" is invoked as the training method of text-davinci-002. The observations of its behavior otherwise stand alone. This is kind of fascinating from an epistemological standpoint. I was
x Mysteries of mode collapse — LessWrong Conjecture (org) RLHF GPT AI Curated 303 Mysteries of mode collapse by janus 8th Nov 2022 AI Alignment Forum 17 min read 57 303 Ω 95 Thanks to Ian McKenzie and Nicholas Dupuis, collaborators on a related project, for contributing to the ideas and experiments discussed in this post. Ian performed some of the random number experiments. Also thanks to Connor Leahy for feedback on a draft, and thanks to Evan Hubinger, Connor Leahy, Beren Millidge, Ethan Perez, Tomek Korbak, Garrett Baker, Leo Gao and various others at Conjecture, Anthropic, and OpenAI for u
Explore this link on the map →related reading
- Mode collapse - Wikipediaen.wikipedia.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- gpt-4.pdfcdn.openai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- models have some pretty funny attractor states — LessWronglesswrong.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Taking LLMs Seriously (As Language Models) — LessWronglesswrong.com
- The N Implementation Details of RLHF with PPO | ICLR Blogposts 2024iclr-blogposts.github.io
- SolidGoldMagikarp (plus, prompt generation) — AI Alignment Forumalignmentforum.org