Interpretability — LessWrong
Chris Olah wrote the following topic prompt for the Open Phil 2021 request for proposals on the alignment of AI systems. We (Asya Bergal and Nick Bec…
x Interpretability — LessWrong Open Philanthropy 2021 AI Alignment RFP Interpretability (ML & AI) AI Frontpage 61 Interpretability by abergal , Nick_Beckstead 29th Oct 2021 AI Alignment Forum 14 min read 13 61 Ω 33 Chris Olah wrote the following topic prompt for the Open Phil 2021 request for proposals on the alignment of AI systems. We (Asya Bergal and Nick Beckstead) are running the Open Phil RFP and are posting each section as a sequence on the Alignment Forum. Although Chris wrote this document, we didn’t want to commit him to being responsible for responding to comments on it by posting i
related reading
- Transformer Circuits Threadtransformer-circuits.pub
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- On Optimism for Interpretabilitygoodfire.ai
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Introduction to Mechanistic Interpretability - by Sarahblog.bluedot.org
- The Building Blocks of Interpretabilitydistill.pub
- Towards Automated Circuit Discovery for Mechanistic Interpretabilityarxiv.org