Rant on Problem Factorization for Alignment — LessWrong
This post is the second in what is likely to become a series of uncharitable rants about alignment proposals (previously: Godzilla Strategies). In general, these posts are intended to convey my underlying intuitions. They are not intended to convey my all-things-considered, reflectively-endorsed opinions. In particular, my all-things-considered reflectively-endorsed opinions are usually more kind. But I think it is valuable to make the underlying, not-particularly-kind intuitions publicly-visible, so people can debate underlying generators directly. I apologize in advance to all the people I insult in the process. With that in mind, let's talk about problem factorization (a.k.a. task decomposition). It all started with HCH, a.k.a. The Infinite Bureaucracy. The idea of The Infinite Bureaucracy is that a human (or, in practice, human-mimicking AI) is given a problem. They only have a small amount of time to think about it and research it, but they can delegate subproblems to their underl
x Rant on Problem Factorization for Alignment — LessWrong "Why Not Just..." Factored Cognition Debate (AI safety technique) Humans consulting HCH Ought AI Frontpage 112 Rant on Problem Factorization for Alignment by johnswentworth 5th Aug 2022 AI Alignment Forum 8 min read 53 112 Ω 45 This post is the second in what is likely to become a series of uncharitable rants about alignment proposals (previously: Godzilla Strategies ). In general, these posts are intended to convey my underlying intuitions. They are not intended to convey my all-things-considered, reflectively-endorsed opinions. In par
related reading
- Rant on Problem Factorization for Alignment — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Tips for Empirical Alignment Research — AI Alignment Forumalignmentforum.org
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequentsequent.org
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- The Field of AI Alignment: A Postmortem, and What To Do About It — LessWronglesswrong.com
- What Is The Alignment Problem? — LessWronglesswrong.com
- A Mike's-Eye View of ARC's Research — Alignment Research Centeralignment.org
- Ngo and Yudkowsky on alignment difficulty — LessWronglesswrong.com
- Automated Alignment is Harder Than You Think — LessWronglesswrong.com