flâneur — a map of the web's best reading

Risks from Learned Optimization: Introduction — AI Alignment Forum

alignmentforum.org · 10,346 words · saved by 1 readers

This is the first of five posts in the Risks from Learned Optimization Sequence based on the paper “Risks from Learned Optimization in Advanced Machine Learning Systems” by Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. Each post in the sequence corresponds to a different section of the paper. Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, and Joar Skalse contributed equally to this sequence. With special thanks to Paul Christiano, Eric Drexler, Rob Bensinger, Jan Leike, Rohin Shah, William Saunders, Buck Shlegeris, David Dalrymple, Abram Demski, Stuart Armstrong, Linda Linsefors, Carl Shulman, Toby Ord, Kate Woolverton, and everyone else who provided feedback on earlier versions of this sequence. The goal of this sequence is to analyze the type of learned optimization that occurs when a learned model (such as a neural network) is itself an optimizer—a situation we refer to as mesa-optimization, a neologism we introduce in this sequence. We

x Risks from Learned Optimization: Introduction — AI Alignment Forum Risks from Learned Optimization Mesa-Optimization Inner Alignment Outer Alignment Optimization AI Risk AI Frontpage 58 Risks from Learned Optimization: Introduction by evhub , Chris van Merwijk , Vlad Mikulik , Joar Skalse , Scott Garrabrant 31st May 2019 14 min read 42 58 This is the first of five posts in the Risks from Learned Optimization Sequence based on the paper “ Risks from Learned Optimization in Advanced Machine Learning Systems ” by Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabr

Explore this link on the map →

related reading