Deceptive Alignment — AI Alignment Forum
This is the fourth of five posts in the Risks from Learned Optimization Sequence based on the paper “Risks from Learned Optimization in Advanced Machine Learning Systems” by Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. Each post in the sequence corresponds to a different section of the paper. With enough training in sufficiently diverse environments, it seems plausible that the base objective will eventually have to be fully represented in the mesa-optimizer. We propose that this can happen without the mesa-optimizer becoming robustly aligned, however. Specifically, a mesa-optimizer might come to model the base objective function and learn that the base optimizer will modify the mesa-optimizer if the mesa-optimizer scores poorly on the base objective. If the mesa-optimizer has an objective that extends across parameter updates, then it will be incentivized to avoid being modified,[1] as it might not pursue the same objective after modification
x Deceptive Alignment — AI Alignment Forum Risks from Learned Optimization Deceptive Alignment AI Risk Mesa-Optimization Instrumental convergence AI Frontpage 39 Deceptive Alignment by evhub , Chris van Merwijk , Vlad Mikulik , Joar Skalse , Scott Garrabrant 5th Jun 2019 20 min read 20 39 This is the fourth of five posts in the Risks from Learned Optimization Sequence based on the paper “ Risks from Learned Optimization in Advanced Machine Learning Systems ” by Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. Each post in the sequence corresponds to a diff
Explore this link on the map →related reading
- Deceptively Aligned Mesa-Optimizers: It's Not Funny If I Have To Explain Itastralcodexten.com
- Risks from Learned Optimization: Introduction — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — AI Alignment Forumalignmentforum.org
- The Inner Alignment Problem — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- How likely is deceptive alignment? — AI Alignment Forumalignmentforum.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Mesa-Optimization — AI Alignment Forumalignmentforum.org
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com