Introduction to Mechanistic Interpretability - by Sarah
Mechanistic Interpretability is an emerging field that seeks to understand the internal reasoning processes of trained neural networks and gain insight into how and why they produce the outputs that they do.
Blog Introduction to Mechanistic Interpretability Sarah Aug 19, 2024 30 3 1 Share Mechanistic Interpretability is an emerging field that seeks to understand the internal reasoning processes of trained neural networks and gain insight into how and why they produce the outputs that they do. AI researchers currently have very little understanding of what is happening inside state-of-the-art models. 1 Current frontier models are extremely large – and extremely complicated. They might contain billions, or even trillions of parameters, spread across over 100 layers. Though we control the data that i
Explore this link on the map →related reading
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Sparsify: A mechanistic interpretability research agenda — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Toy Models of Superpositiontransformer-circuits.pub
- Interpretability — LessWronglesswrong.com
- On Optimism for Interpretabilitygoodfire.ai
- The Building Blocks of Interpretabilitydistill.pub
- Transformer Circuits Threadtransformer-circuits.pub