Mesa-Optimization - AI Alignment Forum
Mesa-Optimization is the situation that occurs when a learned model (such as a neural network) is itself an optimizer. In this situation, a base optimizer creates a second optimizer, called a mesa-optimizer. The primary reference work for this concept is Hubinger et al.'s "Risks from Learned Optimization in Advanced Machine Learning Systems". Example: Natural selection is an optimization process that optimizes for reproductive fitness. Natural selection produced humans, who are themselves optimizers. Humans are therefore mesa-optimizers of natural selection. In the context of AI alignment, the concern is that a base optimizer (e.g., a gradient descent process) may produce a learned model that is itself an optimizer, and that has unexpected and undesirable properties. Even if the gradient descent process is in some sense "trying" to do exactly what human developers want, the resultant mesa-optimizer will not typically be trying to do the exact same thing.[1] Previously work under this c
x Mesa-Optimization — AI Alignment Forum Mesa-Optimization Edited by riceissa , Rob Bensinger , Ruby , et al. last updated 20th Sep 2022 Mesa-Optimization is the situation that occurs when a learned model (such as a neural network) is itself an optimizer. In this situation, a base optimizer creates a second optimizer, called a mesa- optimizer . The primary reference work for this concept is Hubinger et al.'s " Risks from Learned Optimization in Advanced Machine Learning Systems ". Example: Natural selection is an optimization process that optimizes for reproductive fitness. Natural selection p
related reading
- [1906.01820] Risks from Learned Optimization in Advanced Machine Learning Systemsarxiv.org
- Risks from Learned Optimization: Introduction — LessWronglesswrong.com
- Conditions for Mesa-Optimization — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — AI Alignment Forumalignmentforum.org
- Deceptively Aligned Mesa-Optimizers: It's Not Funny If I Have To Explain Itastralcodexten.com
- The Inner Alignment Problem — AI Alignment Forumalignmentforum.org
- [1906.01820] Risks from Learned Optimization in Advanced Machine Learning Systemsarxiv.org
- Deceptive Alignment — AI Alignment Forumalignmentforum.org
- Uncovering mesa-optimization algorithms in Transformersarxiv.org
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com