Better priors as a safety problem — LessWrong
(Related: Inaccessible Information, What does the universal prior actually look like?, Learning the prior) Fitting a neural net implicitly uses a “wrong” prior. This makes neural nets more data hungry and makes them generalize in ways we don’t endorse, but it’s not clear whether it’s an alignment problem. After all, if neural nets are what works, then both the aligned and unaligned AIs will be using them. It’s not clear if that systematically disadvantages aligned AI. Unfortunately I think it’s an alignment problem: In this post I want to try to build some intuition for this problem, and then explain why I’m currently feeling excited about learning the right prior. We usually work with very broad “universal” priors, both in theory (e.g. Solomonoff induction) and in practice (deep neural nets are a very broad hypothesis class). For simplicity I’ll talk about the theoretical setting in this section, but I think the points apply equally well in practice. The classic universal prior is a r
x Better priors as a safety problem — LessWrong AI Rationality Frontpage 67 Better priors as a safety problem by paulfchristiano 5th Jul 2020 ai-alignment.com AI Alignment Forum 6 min read 7 67 Ω 33 ( Related: Inaccessible Information , What does the universal prior actually look like? , Learning the prior ) Fitting a neural net implicitly uses a “wrong” prior. This makes neural nets more data hungry and makes them generalize in ways we don’t endorse, but it’s not clear whether it’s an alignment problem. After all, if neural nets are what works, then both the aligned and unaligned AIs will be
saved by
related reading
- [2601.10160] Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentarxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentalignmentpretraining.ai
- Ilya Sutskever — We're moving from the age of scaling to the age of researchdwarkesh.com
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Bayesian Neural Networkscs.toronto.edu
- Musings on the Speed Prior — AI Alignment Forumalignmentforum.org
- Imitative Generalisation (AKA 'Learning the Prior') — AI Alignment Forumalignmentforum.org
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- The Solomonoff Prior is Malign — LessWronglesswrong.com
- Probabilities are not the right concept — LessWronglesswrong.com