Inner Alignment: Explain like I'm 12 Edition - LessWrong
(This is an unofficial explanation of Inner Alignment based on the Miri paper Risks from Learned Optimization in Advanced Machine Learning Systems (which is almost identical to the LW sequence) and t…
x Inner Alignment: Explain like I'm 12 Edition — LessWrong Best of LessWrong 2020 Inner Alignment Mesa-Optimization AI Frontpage 189 Inner Alignment: Explain like I'm 12 Edition by Rafael Harth 1st Aug 2020 AI Alignment Forum 15 min read 47 189 Ω 59 (This is an unofficial explanation of Inner Alignment based on the Miri paper Risks from Learned Optimization in Advanced Machine Learning Systems (which is almost identical to the LW sequence ) and the Future of Life podcast with Evan Hubinger ( Miri / LW ). It's meant for anyone who found the sequence too long/challenging/technical to read.) Note
Explore this link on the map →related reading
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- The Inner Alignment Problem — AI Alignment Forumalignmentforum.org
- How likely is deceptive alignment? — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Deceptively Aligned Mesa-Optimizers: It's Not Funny If I Have To Explain Itastralcodexten.com
- Does SGD Produce Deceptive Alignment? — LessWronglesswrong.com
- How To Go From Interpretability To Alignment: Just Retarget The Search — LessWronglesswrong.com