Rant on Problem Factorization for Alignment — LessWrong
This post is the second in what is likely to become a series of uncharitable rants about alignment proposals (previously: Godzilla Strategies). In general, these posts are intended to convey my underlying intuitions. They are not intended to convey my all-things-considered, reflectively-endorsed opinions. In particular, my all-things-considered reflectively-endorsed opinions are usually more kind. But I think it is valuable to make the underlying, not-particularly-kind intuitions publicly-visible, so people can debate underlying generators directly. I apologize in advance to all the people I insult in the process. With that in mind, let's talk about problem factorization (a.k.a. task decomposition). It all started with HCH, a.k.a. The Infinite Bureaucracy. The idea of The Infinite Bureaucracy is that a human (or, in practice, human-mimicking AI) is given a problem. They only have a small amount of time to think about it and research it, but they can delegate subproblems to their underl
x Rant on Problem Factorization for Alignment — LessWrong "Why Not Just..." Factored Cognition Debate (AI safety technique) Humans consulting HCH Ought AI Frontpage 112 Rant on Problem Factorization for Alignment by johnswentworth 5th Aug 2022 AI Alignment Forum 8 min read 53 112 Ω 45 This post is the second in what is likely to become a series of uncharitable rants about alignment proposals (previously: Godzilla Strategies ). In general, these posts are intended to convey my underlying intuitions. They are not intended to convey my all-things-considered, reflectively-endorsed opinions. In par
Explore this link on the map →related reading
- Rant on Problem Factorization for Alignment — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- ALIGNMENT - by vincent huang - a slice of my mindmindslice.substack.com
- Tips for Empirical Alignment Research — AI Alignment Forumalignmentforum.org
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequentsequent.org
- The Field of AI Alignment: A Postmortem, and What To Do About It — LessWronglesswrong.com
- What Is The Alignment Problem? — LessWronglesswrong.com
- Ngo and Yudkowsky on alignment difficulty — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com