✳flâneur — a map of the web's best reading
An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers — Neel Nanda
neelnanda.io · 8,668 words · saved by 1 readers
A highly opinionated list of what mechanistic interpretability papers to read when getting into the field
x An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forum Distillation & Pedagogy Interpretability (ML & AI) Sparse Autoencoders (SAEs) Transformer Circuits AI Frontpage 53 An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 by Neel Nanda 7th Jul 2024 30 min read 17 53 This post represents my personal hot takes, not the opinions of my team or employer. This is a massively updated version of a similar list I made two years ago There’s a lot of mechanistic interpretability papers, and more come
Explore this link on the map →related reading
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Transformer Circuits Threadtransformer-circuits.pub
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- Sparsify: A mechanistic interpretability research agenda — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- The Building Blocks of Interpretabilitydistill.pub
- Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] — LessWronglesswrong.com