flâneur — a map of the web's best reading

The Inner Alignment Problem — AI Alignment Forum

alignmentforum.org · 4,692 words · saved by 1 readers

This is the third of five posts in the Risks from Learned Optimization Sequence based on the paper “Risks from Learned Optimization in Advanced Machine Learning Systems” by Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. Each post in the sequence corresponds to a different section of the paper. In this post, we outline reasons to think that a mesa-optimizer may not optimize the same objective function as its base optimizer. Machine learning practitioners have direct control over the base objective function—either by specifying the loss function directly or training a model for it—but cannot directly specify the mesa-objective developed by a mesa-optimizer. We refer to this problem of aligning mesa-optimizers with the base objective as the inner alignment problem. This is distinct from the outer alignment problem, which is the traditional problem of ensuring that the base objective captures the intended goal of the programmers. Current machine lear

x The Inner Alignment Problem — AI Alignment Forum Risks from Learned Optimization Inner Alignment AI Risk Mesa-Optimization AI Frontpage 30 The Inner Alignment Problem by evhub , Chris van Merwijk , Vlad Mikulik , Joar Skalse , Scott Garrabrant 4th Jun 2019 15 min read 17 30 This is the third of five posts in the Risks from Learned Optimization Sequence based on the paper “ Risks from Learned Optimization in Advanced Machine Learning Systems ” by Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. Each post in the sequence corresponds to a different section

Explore this link on the map →

related reading