A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nanda
neelnanda.io · 9,852 words · saved by 2 readers
thank you for existing Blog About Subscribe to hear about new posts (RSS)! Give feedback here!
The main doc is hosted on Dynalist, which has a much better UI for long docs and I highly recommend reading it there. The below is a janky HTML dump for those who prefer this UI Introduction Why does this doc exist? The goal of this doc is to be a comprehensive glossary and explainer for Mechanistic Interpretability (focusing on transformer language models), the field of studying how to reverse engineer neural networks. There's a lot of complex terms and jargon in the field! And these are often scattered across various papers, which tend to be pretty well-written but not designed to be…
saved by
related reading
- Mechanistic Interpretability: Circuits, Induction Headsmbrenndoerfer.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Softmax Linear Unitstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- The Building Blocks of Interpretabilitydistill.pub
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumneelnanda.io