A minimal viable product for alignment - by Jan Leike
Tl;dr: My currently favored approach to solving the alignment problem: automating alignment research using sufficiently aligned AI systems. It doesn’t require humans to solve all alignment problems themselves, and can ultimately help bootstrap better alignment solutions. The space of all problems is truly vast and the space of problems that humanity is currently able to solve is pretty tiny in comparison. This means that today we just aren’t in a position to solve most problems. This is a key motivation to work on AI: progress in AI will significantly expand the space of problems humanity can solve. Maybe a once-and-for-all solution to the alignment problem is located in the space of problems humans can solve. But maybe not. By trying to solve the whole problem, we might be trying to get something that isn’t within our reach. Instead, we can pursue a less ambitious goal that can still ultimately lead us to a solution, a minimal viable product (MVP) for alignment: Building a sufficientl
A minimal viable product for alignment Bootstrapping a solution to the alignment problem Jan Leike Mar 29, 2022 30 Share Tl;dr: My currently favored approach to solving the alignment problem : automating alignment research using sufficiently aligned AI systems. It doesn’t require humans to solve all alignment problems themselves, and can ultimately help bootstrap better alignment solutions. The space of all problems is truly vast and the space of problems that humanity is currently able to solve is pretty tiny in comparison. This means that today we just aren’t in a position to solve most prob
Explore this link on the map →saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- Tips for Empirical Alignment Research — AI Alignment Forumalignmentforum.org
- Why I’m optimistic about our alignment approachaligned.substack.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Why I’m optimistic about our alignment approachaligned.substack.com
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequentsequent.org
- Automated Alignment is Harder Than You Think — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Sequent: scale and automation for higher confidence in alignment — AI Alignment Forumalignmentforum.org
- The Field of AI Alignment: A Postmortem, and What To Do About It — LessWronglesswrong.com
- Short Timelines Don't Devalue Long Horizon Research — LessWronglesswrong.com