The Plan - 2023 Version — LessWrong
Background: The Plan, The Plan: 2022 Update. If you haven’t read those, don’t worry, we’re going to go through things from the top this year, and wit…
x The Plan - 2023 Version — LessWrong Agent Foundations AI Risk Meta-Philosophy Natural Abstraction Research Agendas AI Frontpage 153 The Plan - 2023 Version by johnswentworth 29th Dec 2023 38 min read 41 153 Background: The Plan , The Plan: 2022 Update . If you haven’t read those, don’t worry, we’re going to go through things from the top this year, and with moderately more detail than before. 1. What’s Your Plan For AI Alignment? Median happy trajectory: Sort out our fundamental confusions about agency and abstraction enough to do interpretability that works and generalizes robustly. Look th
Explore this link on the map →related reading
- The Plan - 2025 Update — LessWronglesswrong.com
- LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- 2023 letter | Zhengdongzhengdongwang.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Shah and Yudkowsky on alignment failures — LessWronglesswrong.com
- On how various plans miss the hard bits of the alignment challenge — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Why Agent Foundations? An Overly Abstract Explanation — LessWronglesswrong.com
- Dialogue: Is there a Natural Abstraction of Good? — LessWronglesswrong.com
- The Best of LessWrong — LessWronglesswrong.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org