How To Go From Interpretability To Alignment: Just Retarget The Search — LessWrong
Here's a simple strategy for AI alignment: use interpretability tools to identify the AI's internal search process, and the AI's internal representat…
x How To Go From Interpretability To Alignment: Just Retarget The Search — LessWrong "Why Not Just..." Inner Alignment Interpretability (ML & AI) AI Risk AI Frontpage 214 How To Go From Interpretability To Alignment: Just Retarget The Search by johnswentworth 10th Aug 2022 AI Alignment Forum 3 min read 34 214 Ω 76 [EDIT: Many people who read this post were very confused about some things, which I later explained in What’s General-Purpose Search, And Why Might We Expect To See It In Trained ML Systems? You might want to read that post first.] When people talk about prosaic alignment proposals ,
Explore this link on the map →saved by
related reading
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Prospects for Alignment Automation: Interpretability Case Study — LessWronglesswrong.com
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- (My understanding of) What Everyone in Technical Alignment is Doing and Why — LessWronglesswrong.com
- Import AIjack-clark.net
- Stream of Search (SoS): Learning to Search in Languagearxiv.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com