Low-stakes alignment — AI Alignment Forum
Right now I’m working on finding a good objective to optimize with ML, rather than trying to make sure our models are robustly optimizing that objective. (This is roughly “outer alignment.”) That’s pretty vague, and it’s not obvious whether “find a good objective” is a meaningful goal rather than being inherently confused or sweeping key distinctions under the rug. So I like to focus on a more precise special case of alignment: solve alignment when decisions are “low stakes.” I think this case effectively isolates the problem of “find a good objective” from the problem of ensuring robustness and is precise enough to focus on productively. In this post I’ll describe what I mean by the low-stakes setting, why I think it isolates this subproblem, why I want to isolate this subproblem, and why I think that it’s valuable to work on crisp subproblems. A situation is low-stakes if we care very little about any small number of decisions. That is, we only care about the average behavior of the
x Low-stakes alignment — AI Alignment Forum AI Frontpage 45 Low-stakes alignment by paulfchristiano 30th Apr 2021 ai-alignment.com 8 min read 11 45 Right now I’m working on finding a good objective to optimize with ML, rather than trying to make sure our models are robustly optimizing that objective. (This is roughly “ outer alignment .”) That’s pretty vague, and it’s not obvious whether “find a good objective” is a meaningful goal rather than being inherently confused or sweeping key distinctions under the rug. So I like to focus on a more precise special case of alignment: solve alignment wh
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Outer vs inner misalignment: three framings — AI Alignment Forumalignmentforum.org
- Why I’m optimistic about our alignment approachaligned.substack.com
- Outer vs inner misalignment: three framings — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Why AI alignment could be hard with modern deep learningcold-takes.com