flâneur — a map of the web's best reading

Inner Alignment: Explain like I'm 12 Edition — LessWrong

lesswrong.com · 9,221 words · saved by 1 readers

(This is an unofficial explanation of Inner Alignment based on the Miri paper Risks from Learned Optimization in Advanced Machine Learning Systems (which is almost identical to the LW sequence) and the Future of Life podcast with Evan Hubinger (Miri/LW). It's meant for anyone who found the sequence too long/challenging/technical to read.) Note that bold and italics means "this is a new term I'm introducing," whereas underline and italics is used for emphasis. Let's start with an abridged guide to how Deep Learning works: If the problem is "find a tool that can look at any image and decide whether or not it contains a cat," then each conceivable set of rules for answering this question (formally, each function from the set of all pixels to the set { yes , no } ) defines one solution. We call each such solution a model. The space of possible models is depicted below. Since that's all possible models, most of them are utter nonsense. Pick a random one, and you're as likely to end up with

x Inner Alignment: Explain like I'm 12 Edition — LessWrong Alignment & Agency Inner Alignment Mesa-Optimization AI Frontpage 188 Inner Alignment: Explain like I'm 12 Edition by Rafael Harth 1st Aug 2020 AI Alignment Forum 15 min read 47 188 Ω 59 (This is an unofficial explanation of Inner Alignment based on the Miri paper Risks from Learned Optimization in Advanced Machine Learning Systems (which is almost identical to the LW sequence ) and the Future of Life podcast with Evan Hubinger ( Miri / LW ). It's meant for anyone who found the sequence too long/challenging/technical to read.) Note tha

Explore this link on the map →

saved by

related reading