Alignment By Default — AI Alignment Forum
Suppose AI continues on its current trajectory: deep learning continues to get better as we throw more data and compute at it, researchers keep trying random architectures and using whatever seems to work well in practice. Do we end up with aligned AI “by default”? I think there’s at least a plausible trajectory in which the answer is “yes”. Not very likely - I’d put it at ~10% chance - but plausible. In fact, there’s at least an argument to be made that alignment-by-default is more likely to work than many fancy alignment proposals, including IRL variants and HCH-family methods. This post presents the rough models and arguments. I’ll break it down into two main pieces: Ultimately, we’ll consider a semi-supervised/transfer-learning style approach, where we first do some unsupervised learning and hopefully “learn human values” before starting the supervised/reinforcement part. As background, I will assume you’ve read some of the core material about human values from the sequences, inclu
x Alignment By Default — AI Alignment Forum Best of LessWrong 2020 Abstraction Natural Abstraction AI Frontpage 64 Alignment By Default by johnswentworth 12th Aug 2020 14 min read 101 64 Suppose AI continues on its current trajectory: deep learning continues to get better as we throw more data and compute at it, researchers keep trying random architectures and using whatever seems to work well in practice. Do we end up with aligned AI “by default”? I think there’s at least a plausible trajectory in which the answer is “yes”. Not very likely - I’d put it at ~10% chance - but plausible. In fact,
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- The Artificiality of Alignmentjoinreboot.org
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Alignment Is Proven To Be Solvable - by SE Gygesverysane.ai
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- Why AI alignment could be hard with modern deep learningcold-takes.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Value Alignment Is a Pseudo Conceptzilanqian.substack.com
- What Is The Alignment Problem? — LessWronglesswrong.com
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org