The Learning-Theoretic Agenda: Status 2023 — AI Alignment Forum
TLDR: I give an overview of (i) the key problems that the learning-theoretic AI alignment research agenda is trying to solve, (ii) the insights we have so far about these problems and (iii) the research directions I know for attacking these problems. I also describe "Physicalist Superimitation" (previously knows as "PreDCA"): a hypothesized alignment protocol based on infra-Bayesian physicalism that is designed to reliably learn and act on the user's values. I wish to thank Steven Byrnes, Abram Demski, Daniel Filan, Felix Harder, Alexander Gietelink Oldenziel, John S. Wentworth[1] and my spouse Marcus Ogren for reviewing a draft of this article, finding errors and making helpful comments and suggestions. Any remaining flaws are entirely my fault. I already described the learning-theoretic agenda (henceforth LTA) in 2018. While the overall spirit stayed similar, naturally LTA has evolved since then, with new results and research directions, and slightly different priorities. This calls
x The Learning-Theoretic Agenda: Status 2023 — AI Alignment Forum window.__lwSsrGql.inject("query postCommentsThreadQuery($selector: CommentSelector, $limit: Int, $enableTotal: Boolean) {\n comments(selector: $selector, limit: $limit, enableTotal: $enableTotal) {\n results {\n ...CommentsList\n }\n totalCount\n }\n}\n\nfragment TagBasicInfo on Tag {\n _id\n userId\n name\n shortName\n slug\n core\n postCount\n adminOnly\n canEditUserIds\n suggestedAsFilter\n needsReview\n descriptionTruncationCount\n createdAt\n wikiOnly\n deleted\n isSubforum\n noindex\n isArbitalImport\n isPlaceholderPage\n
Explore this link on the map →related reading
- LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Study Guide — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- ALIGNMENT - by vincent huang - a slice of my mindmindslice.substack.com
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Shtetl-Optimized >> Blog Archive >> Theory and AI Alignmentscottaaronson.blog
- The Field of AI Alignment: A Postmortem, and What To Do About It — LessWronglesswrong.com
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequentsequent.org
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- The Best of LessWrong — LessWronglesswrong.com