The Plan - 2023 Version — LessWrong
Background: The Plan, The Plan: 2022 Update. If you haven’t read those, don’t worry, we’re going to go through things from the top this year, and wit…
x The Plan - 2023 Version — LessWrong Agent Foundations AI Risk Meta-Philosophy Natural Abstraction Research Agendas AI Frontpage 153 The Plan - 2023 Version by johnswentworth 29th Dec 2023 38 min read 41 153 Background: The Plan , The Plan: 2022 Update . If you haven’t read those, don’t worry, we’re going to go through things from the top this year, and with moderately more detail than before. 1. What’s Your Plan For AI Alignment? Median happy trajectory: Sort out our fundamental confusions about agency and abstraction enough to do interpretability that works and generalizes robustly. Look th
related reading
- The Plan - 2025 Update — LessWronglesswrong.com
- LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- 2023 letter | Zhengdongzhengdongwang.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Shah and Yudkowsky on alignment failures — LessWronglesswrong.com
- On how various plans miss the hard bits of the alignment challenge — LessWronglesswrong.com
- The Best of LessWrong — LessWronglesswrong.com
- Why Agent Foundations? An Overly Abstract Explanation — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Dialogue: Is there a Natural Abstraction of Good? — LessWronglesswrong.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org