flâneur — a map of the web's best reading

Deceptive Alignment — AI Alignment Forum

alignmentforum.org · 7,367 words · saved by 1 readers

This is the fourth of five posts in the Risks from Learned Optimization Sequence based on the paper “Risks from Learned Optimization in Advanced Machine Learning Systems” by Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. Each post in the sequence corresponds to a different section of the paper. With enough training in sufficiently diverse environments, it seems plausible that the base objective will eventually have to be fully represented in the mesa-optimizer. We propose that this can happen without the mesa-optimizer becoming robustly aligned, however. Specifically, a mesa-optimizer might come to model the base objective function and learn that the base optimizer will modify the mesa-optimizer if the mesa-optimizer scores poorly on the base objective. If the mesa-optimizer has an objective that extends across parameter updates, then it will be incentivized to avoid being modified,[1] as it might not pursue the same objective after modification

x Deceptive Alignment — AI Alignment Forum Risks from Learned Optimization Deceptive Alignment AI Risk Mesa-Optimization Instrumental convergence AI Frontpage 39 Deceptive Alignment by evhub , Chris van Merwijk , Vlad Mikulik , Joar Skalse , Scott Garrabrant 5th Jun 2019 20 min read 20 39 This is the fourth of five posts in the Risks from Learned Optimization Sequence based on the paper “ Risks from Learned Optimization in Advanced Machine Learning Systems ” by Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. Each post in the sequence corresponds to a diff

Explore this link on the map →

related reading