✳flâneur — a map of the web's best reading
Outer vs inner misalignment: three framings — LessWrong
lesswrong.com · 4,040 words · saved by 1 readers
A core concept in the field of AI alignment is a distinction between two types of misalignment: outer misalignment and inner misalignment. Roughly sp…
x Outer vs inner misalignment: three framings — LessWrong Inner Alignment Outer Alignment AI Frontpage 53 Outer vs inner misalignment: three framings by Richard_Ngo 6th Jul 2022 AI Alignment Forum 11 min read 5 53 Ω 28 A core concept in the field of AI alignment is a distinction between two types of misalignment: outer misalignment and inner misalignment. Roughly speaking, the outer alignment problem is the problem of specifying an reward function which captures human preferences; and the inner alignment problem is the problem of ensuring that a policy trained on that reward function actually
Explore this link on the map →related reading
- Outer vs inner misalignment: three framings — AI Alignment Forumalignmentforum.org
- Categorizing failures as “outer” or “inner” misalignment is often confused — AI Alignment Forumalignmentforum.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org
- What is AI alignment? - by Adam Jones - BlueDot Impactaisafetyfundamentals.com
- What is AI alignment? - by Adam Jones - BlueDot Impactbluedot.org
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Low-stakes alignment — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org