Mesa-Optimization - AI Alignment Forum
Mesa-Optimization is the situation that occurs when a learned model (such as a neural network) is itself an optimizer. In this situation, a base optimizer creates a second optimizer, called a mesa-optimizer. The primary reference work for this concept is Hubinger et al.'s "Risks from Learned Optimization in Advanced Machine Learning Systems". Example: Natural selection is an optimization process that optimizes for reproductive fitness. Natural selection produced humans, who are themselves optimizers. Humans are therefore mesa-optimizers of natural selection. In the context of AI alignment, the concern is that a base optimizer (e.g., a gradient descent process) may produce a learned model that is itself an optimizer, and that has unexpected and undesirable properties. Even if the gradient descent process is in some sense "trying" to do exactly what human developers want, the resultant mesa-optimizer will not typically be trying to do the exact same thing.[1] Previously work under this c
x Mesa-Optimization — AI Alignment Forum Mesa-Optimization Edited by riceissa , Rob Bensinger , Ruby , et al. last updated 20th Sep 2022 Mesa-Optimization is the situation that occurs when a learned model (such as a neural network) is itself an optimizer. In this situation, a base optimizer creates a second optimizer, called a mesa- optimizer . The primary reference work for this concept is Hubinger et al.'s " Risks from Learned Optimization in Advanced Machine Learning Systems ". Example: Natural selection is an optimization process that optimizes for reproductive fitness. Natural selection p
Explore this link on the map →related reading
- Conditions for Mesa-Optimization — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — LessWronglesswrong.com
- [1906.01820] Risks from Learned Optimization in Advanced Machine Learning Systemsarxiv.org
- Deceptively Aligned Mesa-Optimizers: It's Not Funny If I Have To Explain Itastralcodexten.com
- The Inner Alignment Problem — AI Alignment Forumalignmentforum.org
- [1906.01820] Risks from Learned Optimization in Advanced Machine Learning Systemsarxiv.org
- Deceptive Alignment — AI Alignment Forumalignmentforum.org
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com