flâneur — a map of the web's best reading

The case for aligning narrowly superhuman models — LessWrong

lesswrong.com · 21,458 words · saved by 1 readers

I wrote this post to get people’s takes on a type of work that seems exciting to me personally; I’m not speaking for Open Phil as a whole. Institutionally, we are very uncertain whether to prioritize this (and if we do where it should be housed and how our giving should be structured). We are not seeking grant applications on this topic right now. Thanks to Daniel Dewey, Eliezer Yudkowsky, Evan Hubinger, Holden Karnofsky, Jared Kaplan, Mike Levine, Nick Beckstead, Owen Cotton-Barratt, Paul Christiano, Rob Bensinger, and Rohin Shah for comments on earlier drafts. A genre of technical AI risk reduction work that seems exciting to me is trying to align existing models that already are, or have the potential to be, “superhuman”[1] at some particular task (which I’ll call narrowly superhuman models).[2] I don’t just mean “train these models to be more robust, reliable, interpretable, etc” (though that seems good too); I mean “figure out how to harness their full abilities so they can be as

x The case for aligning narrowly superhuman models — LessWrong GPT Language Models (LLMs) Machine Learning (ML) Outer Alignment AI Curated 186 The case for aligning narrowly superhuman models by Ajeya Cotra 5th Mar 2021 AI Alignment Forum 46 min read 75 186 Ω 74 I wrote this post to get people’s takes on a type of work that seems exciting to me personally; I’m not speaking for Open Phil as a whole. Institutionally, we are very uncertain whether to prioritize this (and if we do where it should be housed and how our giving should be structured). We are not seeking grant applications on this topi

Explore this link on the map →

related reading