The case for aligning narrowly superhuman models — LessWrong
I wrote this post to get people’s takes on a type of work that seems exciting to me personally; I’m not speaking for Open Phil as a whole. Institutionally, we are very uncertain whether to prioritize this (and if we do where it should be housed and how our giving should be structured). We are not seeking grant applications on this topic right now. Thanks to Daniel Dewey, Eliezer Yudkowsky, Evan Hubinger, Holden Karnofsky, Jared Kaplan, Mike Levine, Nick Beckstead, Owen Cotton-Barratt, Paul Christiano, Rob Bensinger, and Rohin Shah for comments on earlier drafts. A genre of technical AI risk reduction work that seems exciting to me is trying to align existing models that already are, or have the potential to be, “superhuman”[1] at some particular task (which I’ll call narrowly superhuman models).[2] I don’t just mean “train these models to be more robust, reliable, interpretable, etc” (though that seems good too); I mean “figure out how to harness their full abilities so they can be as
x The case for aligning narrowly superhuman models — LessWrong GPT Language Models (LLMs) Machine Learning (ML) Outer Alignment AI Curated 186 The case for aligning narrowly superhuman models by Ajeya Cotra 5th Mar 2021 AI Alignment Forum 46 min read 75 186 Ω 74 I wrote this post to get people’s takes on a type of work that seems exciting to me personally; I’m not speaking for Open Phil as a whole. Institutionally, we are very uncertain whether to prioritize this (and if we do where it should be housed and how our giving should be structured). We are not seeking grant applications on this topi
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- LessWronglesswrong.com
- ALIGNMENT - by vincent huang - a slice of my mindmindslice.substack.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- PSA: Almost nobody is directly working on superintelligent alignment — LessWronglesswrong.com
- OpenAI Launches Superalignment Taskforcethezvi.substack.com
- Why I’m optimistic about our alignment approachaligned.substack.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequentsequent.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com