Thoughts on the Alignment Implications of Scaling Language Models | Leo Gao
[Epistemic status: slightly rambly, mostly personal intuition and opinion that will probably be experimentally proven wrong within a year considering how fast stuff moves in this field] Thanks to Gwern Branwen, Steven Byrnes, Dan Hendrycks, Connor Leahy, Adam Shimi, Kyle and Laria for the insightful discussions and feedback. By now, most of you have probably heard about GPT-3 and what it does. There’s been a bunch of different opinions on what it means for alignment, and this post is yet another opinion from a slightly different perspective. Some background: I’m a part of EleutherAI, a decentralized research collective (read: glorified discord server - come join us on Discord for ML, alignment, and dank memes). We’re best known for our ongoing effort to create a GPT-3-like large language model, and so we have a lot of experience working with transformer models and looking at scaling laws, but we also take alignment very seriously and spend a lot of time thinking about it. (see here for
Created: 2021-06-02   Modified: 2021-06-02 [Epistemic status: slightly rambly, mostly personal intuition and opinion that will probably be experimentally proven wrong within a year considering how fast stuff moves in this field] Thanks to Gwern Branwen, Steven Byrnes, Dan Hendrycks, Connor Leahy, Adam Shimi, Kyle and Laria for the insightful discussions and feedback. Background By now, most of you have probably heard about GPT-3 and what it does. There’s been a bunch of different opinions on what it means for alignment, and this post is yet another opinion from a slightly different perspe
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- larger language models may disappoint you [or, an eternally unfinished draft] — LessWronglesswrong.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- Language Models, World Models, and Human Model-Buildinglingo.csail.mit.edu
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- I am worried about near-term non-LLM AI developments — LessWronglesswrong.com
- chinchilla's wild implications — AI Alignment Forumalignmentforum.org
- Why I’m optimistic about our alignment approachaligned.substack.com
- Announcing the Inverse Scaling Prize ($250k Prize Pool) — AI Alignment Forumalignmentforum.org