flâneur — a map of the web's best reading

The Learning-Theoretic Agenda: Status 2023 — AI Alignment Forum

alignmentforum.org · 1,456 words · saved by 1 readers

TLDR: I give an overview of (i) the key problems that the learning-theoretic AI alignment research agenda is trying to solve, (ii) the insights we have so far about these problems and (iii) the research directions I know for attacking these problems. I also describe "Physicalist Superimitation" (previously knows as "PreDCA"): a hypothesized alignment protocol based on infra-Bayesian physicalism that is designed to reliably learn and act on the user's values. I wish to thank Steven Byrnes, Abram Demski, Daniel Filan, Felix Harder, Alexander Gietelink Oldenziel, John S. Wentworth[1] and my spouse Marcus Ogren for reviewing a draft of this article, finding errors and making helpful comments and suggestions. Any remaining flaws are entirely my fault. I already described the learning-theoretic agenda (henceforth LTA) in 2018. While the overall spirit stayed similar, naturally LTA has evolved since then, with new results and research directions, and slightly different priorities. This calls

x The Learning-Theoretic Agenda: Status 2023 — AI Alignment Forum window.__lwSsrGql.inject("query postCommentsThreadQuery($selector: CommentSelector, $limit: Int, $enableTotal: Boolean) {\n comments(selector: $selector, limit: $limit, enableTotal: $enableTotal) {\n results {\n ...CommentsList\n }\n totalCount\n }\n}\n\nfragment TagBasicInfo on Tag {\n _id\n userId\n name\n shortName\n slug\n core\n postCount\n adminOnly\n canEditUserIds\n suggestedAsFilter\n needsReview\n descriptionTruncationCount\n createdAt\n wikiOnly\n deleted\n isSubforum\n noindex\n isArbitalImport\n isPlaceholderPage\n

Explore this link on the map →

related reading