Conditions for Mesa-Optimization — AI Alignment Forum
This is the second of five posts in the Risks from Learned Optimization Sequence based on the paper “Risks from Learned Optimization in Advanced Machine Learning Systems” by Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. Each post in the sequence corresponds to a different section of the paper. In this post, we consider how the following two components of a particular machine learning system might influence whether it will produce a mesa-optimizer: We deliberately choose to present theoretical considerations for why mesa-optimization may or may not occur rather than provide concrete examples. Mesa-optimization is a phenomenon that we believe will occur mainly in machine learning systems that are more advanced than those that exist today.[1] Thus, an attempt to induce mesa-optimization in a current machine learning system would likely require us to use an artificial setup specifically designed to induce mesa-optimization. Moreover, the limited int
x Conditions for Mesa-Optimization — AI Alignment Forum Risks from Learned Optimization Mesa-Optimization AI Risk AI Frontpage 30 Conditions for Mesa-Optimization by evhub , Chris van Merwijk , Vlad Mikulik , Joar Skalse , Scott Garrabrant 1st Jun 2019 14 min read 48 30 This is the second of five posts in the Risks from Learned Optimization Sequence based on the paper “ Risks from Learned Optimization in Advanced Machine Learning Systems ” by Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. Each post in the sequence corresponds to a different section of th
Explore this link on the map →related reading
- Risks from Learned Optimization: Introduction — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — AI Alignment Forumalignmentforum.org
- Mesa-Optimization — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — LessWronglesswrong.com
- The Inner Alignment Problem — AI Alignment Forumalignmentforum.org
- [1906.01820] Risks from Learned Optimization in Advanced Machine Learning Systemsarxiv.org
- [1906.01820] Risks from Learned Optimization in Advanced Machine Learning Systemsarxiv.org
- Deceptively Aligned Mesa-Optimizers: It's Not Funny If I Have To Explain Itastralcodexten.com
- Deceptive Alignment — AI Alignment Forumalignmentforum.org
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com