flâneur — a map of the web's best reading

Inner Alignment: Explain like I'm 12 Edition - LessWrong

lesswrong.com · 9,197 words · saved by 1 readers

(This is an unofficial explanation of Inner Alignment based on the Miri paper Risks from Learned Optimization in Advanced Machine Learning Systems (which is almost identical to the LW sequence) and t…

x Inner Alignment: Explain like I'm 12 Edition — LessWrong Best of LessWrong 2020 Inner Alignment Mesa-Optimization AI Frontpage 189 Inner Alignment: Explain like I'm 12 Edition by Rafael Harth 1st Aug 2020 AI Alignment Forum 15 min read 47 189 Ω 59 (This is an unofficial explanation of Inner Alignment based on the Miri paper Risks from Learned Optimization in Advanced Machine Learning Systems (which is almost identical to the LW sequence ) and the Future of Life podcast with Evan Hubinger ( Miri / LW ). It's meant for anyone who found the sequence too long/challenging/technical to read.) Note

Explore this link on the map →

related reading