Interpretability — LessWrong
Chris Olah wrote the following topic prompt for the Open Phil 2021 request for proposals on the alignment of AI systems. We (Asya Bergal and Nick Bec…
x Interpretability — LessWrong Open Philanthropy 2021 AI Alignment RFP Interpretability (ML & AI) AI Frontpage 61 Interpretability by abergal , Nick_Beckstead 29th Oct 2021 AI Alignment Forum 14 min read 13 61 Ω 33 Chris Olah wrote the following topic prompt for the Open Phil 2021 request for proposals on the alignment of AI systems. We (Asya Bergal and Nick Beckstead) are running the Open Phil RFP and are posting each section as a sequence on the Alignment Forum. Although Chris wrote this document, we didn’t want to commit him to being responsible for responding to comments on it by posting i
Explore this link on the map →related reading
- Transformer Circuits Threadtransformer-circuits.pub
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- On Optimism for Interpretabilitygoodfire.ai
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- Introduction to Mechanistic Interpretability - by Sarahblog.bluedot.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- The Building Blocks of Interpretabilitydistill.pub
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- Softmax Linear Unitstransformer-circuits.pub