Better priors as a safety problem — LessWrong
(Related: Inaccessible Information, What does the universal prior actually look like?, Learning the prior) Fitting a neural net implicitly uses a “wrong” prior. This makes neural nets more data hungry and makes them generalize in ways we don’t endorse, but it’s not clear whether it’s an alignment problem. After all, if neural nets are what works, then both the aligned and unaligned AIs will be using them. It’s not clear if that systematically disadvantages aligned AI. Unfortunately I think it’s an alignment problem: In this post I want to try to build some intuition for this problem, and then explain why I’m currently feeling excited about learning the right prior. We usually work with very broad “universal” priors, both in theory (e.g. Solomonoff induction) and in practice (deep neural nets are a very broad hypothesis class). For simplicity I’ll talk about the theoretical setting in this section, but I think the points apply equally well in practice. The classic universal prior is a r
x Better priors as a safety problem — LessWrong AI Rationality Frontpage 67 Better priors as a safety problem by paulfchristiano 5th Jul 2020 ai-alignment.com AI Alignment Forum 6 min read 7 67 Ω 33 ( Related: Inaccessible Information , What does the universal prior actually look like? , Learning the prior ) Fitting a neural net implicitly uses a “wrong” prior. This makes neural nets more data hungry and makes them generalize in ways we don’t endorse, but it’s not clear whether it’s an alignment problem. After all, if neural nets are what works, then both the aligned and unaligned AIs will be
Explore this link on the map →saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Musings on the Speed Prior — AI Alignment Forumalignmentforum.org
- Imitative Generalisation (AKA 'Learning the Prior') — AI Alignment Forumalignmentforum.org
- Bayesian Neural Networkscs.toronto.edu
- The Solomonoff Prior is Malign — LessWronglesswrong.com
- Probabilities are not the right concept — LessWronglesswrong.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- What does the universal prior actually look like? | Ordinary Ideasordinaryideas.wordpress.com
- Unfalsifiable stories of doom | Mechanize, Inc.mechanize.work
- Alignment By Default — AI Alignment Forumalignmentforum.org