Risks from Learned Optimization: Introduction — AI Alignment Forum
This is the first of five posts in the Risks from Learned Optimization Sequence based on the paper “Risks from Learned Optimization in Advanced Machine Learning Systems” by Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. Each post in the sequence corresponds to a different section of the paper. Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, and Joar Skalse contributed equally to this sequence. With special thanks to Paul Christiano, Eric Drexler, Rob Bensinger, Jan Leike, Rohin Shah, William Saunders, Buck Shlegeris, David Dalrymple, Abram Demski, Stuart Armstrong, Linda Linsefors, Carl Shulman, Toby Ord, Kate Woolverton, and everyone else who provided feedback on earlier versions of this sequence. The goal of this sequence is to analyze the type of learned optimization that occurs when a learned model (such as a neural network) is itself an optimizer—a situation we refer to as mesa-optimization, a neologism we introduce in this sequence. We
x Risks from Learned Optimization: Introduction — AI Alignment Forum Risks from Learned Optimization Mesa-Optimization Inner Alignment Outer Alignment Optimization AI Risk AI Frontpage 58 Risks from Learned Optimization: Introduction by evhub , Chris van Merwijk , Vlad Mikulik , Joar Skalse , Scott Garrabrant 31st May 2019 14 min read 42 58 This is the first of five posts in the Risks from Learned Optimization Sequence based on the paper “ Risks from Learned Optimization in Advanced Machine Learning Systems ” by Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabr
Explore this link on the map →related reading
- Risks from Learned Optimization: Introduction — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — LessWronglesswrong.com
- The Inner Alignment Problem — AI Alignment Forumalignmentforum.org
- Conditions for Mesa-Optimization — AI Alignment Forumalignmentforum.org
- Deceptive Alignment — AI Alignment Forumalignmentforum.org
- Mesa-Optimization — AI Alignment Forumalignmentforum.org
- [1906.01820] Risks from Learned Optimization in Advanced Machine Learning Systemsarxiv.org
- [1906.01820] Risks from Learned Optimization in Advanced Machine Learning Systemsarxiv.org
- Deceptively Aligned Mesa-Optimizers: It's Not Funny If I Have To Explain Itastralcodexten.com
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com