✳flâneur — a map of the web's best reading
Deceptively Aligned Mesa-Optimizers: It's Not Funny If I Have To Explain It
astralcodexten.com · 3,240 words · saved by 1 readers
A Machine Alignment Monday post, 4/11/22
Deceptively Aligned Mesa-Optimizers: It's Not Funny If I Have To Explain It A Machine Alignment Monday post, 4/11/22 Scott Alexander Apr 11, 2022 117 311 2 Share I. Our goal here is to popularize obscure and hard-to-understand areas of AI alignment, and surely this meme (retweeted by Eliezer last week) qualifies: Leo Gao @nabla_theta 4:24 AM · Dec 13, 2021 2 Reposts · 42 Likes So let’s try to understand the incomprehensible meme! Our main source will be Hubinger et al 2019, Risks From Learned Optimization In Advanced Machine Learning Systems . Mesa- is a Greek prefix which means the opposite o
Explore this link on the map →related reading
- Deceptive Alignment — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — LessWronglesswrong.com
- Risks from Learned Optimization: Introduction — AI Alignment Forumalignmentforum.org
- The Inner Alignment Problem — AI Alignment Forumalignmentforum.org
- Risks from Learned Optimization: Introduction — AI Alignment Forumalignmentforum.org
- Mesa-Optimization — AI Alignment Forumalignmentforum.org
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Conditions for Mesa-Optimization — AI Alignment Forumalignmentforum.org
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org