Alignment By Default — AI Alignment Forum
Suppose AI continues on its current trajectory: deep learning continues to get better as we throw more data and compute at it, researchers keep trying random architectures and using whatever seems to work well in practice. Do we end up with aligned AI “by default”? I think there’s at least a plausible trajectory in which the answer is “yes”. Not very likely - I’d put it at ~10% chance - but plausible. In fact, there’s at least an argument to be made that alignment-by-default is more likely to work than many fancy alignment proposals, including IRL variants and HCH-family methods. This post presents the rough models and arguments. I’ll break it down into two main pieces: Ultimately, we’ll consider a semi-supervised/transfer-learning style approach, where we first do some unsupervised learning and hopefully “learn human values” before starting the supervised/reinforcement part. As background, I will assume you’ve read some of the core material about human values from the sequences, inclu
x Alignment By Default — AI Alignment Forum Best of LessWrong 2020 Abstraction Natural Abstraction AI Frontpage 64 Alignment By Default by johnswentworth 12th Aug 2020 14 min read 101 64 Suppose AI continues on its current trajectory: deep learning continues to get better as we throw more data and compute at it, researchers keep trying random architectures and using whatever seems to work well in practice. Do we end up with aligned AI “by default”? I think there’s at least a plausible trajectory in which the answer is “yes”. Not very likely - I’d put it at ~10% chance - but plausible. In fact,
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Why AI alignment could be hard with modern deep learningcold-takes.com
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- What Is The Alignment Problem? — LessWronglesswrong.com
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Shah and Yudkowsky on alignment failures — LessWronglesswrong.com
- Another (outer) alignment failure story — AI Alignment Forumalignmentforum.org
- Alignment Is Proven To Be Solvable - by SE Gygesverysane.ai