flâneur — a map of the web's best reading

Low-stakes alignment — AI Alignment Forum

alignmentforum.org · 4,265 words · saved by 1 readers

Right now I’m working on finding a good objective to optimize with ML, rather than trying to make sure our models are robustly optimizing that objective. (This is roughly “outer alignment.”) That’s pretty vague, and it’s not obvious whether “find a good objective” is a meaningful goal rather than being inherently confused or sweeping key distinctions under the rug. So I like to focus on a more precise special case of alignment: solve alignment when decisions are “low stakes.” I think this case effectively isolates the problem of “find a good objective” from the problem of ensuring robustness and is precise enough to focus on productively. In this post I’ll describe what I mean by the low-stakes setting, why I think it isolates this subproblem, why I want to isolate this subproblem, and why I think that it’s valuable to work on crisp subproblems. A situation is low-stakes if we care very little about any small number of decisions. That is, we only care about the average behavior of the

x Low-stakes alignment — AI Alignment Forum AI Frontpage 45 Low-stakes alignment by paulfchristiano 30th Apr 2021 ai-alignment.com 8 min read 11 45 Right now I’m working on finding a good objective to optimize with ML, rather than trying to make sure our models are robustly optimizing that objective. (This is roughly “ outer alignment .”) That’s pretty vague, and it’s not obvious whether “find a good objective” is a meaningful goal rather than being inherently confused or sweeping key distinctions under the rug. So I like to focus on a more precise special case of alignment: solve alignment wh

Explore this link on the map →

related reading