✳flâneur — a map of the web's best reading
EIS III: Broad Critiques of Interpretability Research — AI Alignment Forum
alignmentforum.org · 3,590 words · saved by 1 readers
Part 3 of 12 in the Engineer’s Interpretability Sequence. …
x EIS III: Broad Critiques of Interpretability Research — AI Alignment Forum The Engineer’s Interpretability Sequence Interpretability (ML & AI) Research Agendas AI Frontpage 8 EIS III: Broad Critiques of Interpretability Research by scasper 14th Feb 2023 13 min read 2 8 Part 3 of 12 in the Engineer’s Interpretability Sequence . Right now, interpretability is a major subfield in the machine learning research community. As mentioned in EIS I, there is so much work in interpretability that there is now a database of 5199 interpretability papers (Jacovi, 2023) . You can also look at a survey from
Explore this link on the map →related reading
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- On Optimism for Interpretabilitygoodfire.ai
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- Martian Interpretability Challenge: The Core Problems In Interpretability — LessWronglesswrong.com
- The Building Blocks of Interpretabilitydistill.pub
- EIS XIV: Is mechanistic interpretability about to be practically useful? — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com