Mechanistic Interpretability, Variables, and the Importance of Interpretable Bases
Mechanistic interpretability seeks to reverse engineer neural networks, similar to how one might reverse engineer a compiled binary computer program. After all, neural network parameters are in some sense a binary computer program which runs on one of the exotic virtual machines we call a neural network architecture. This is actually quite a deep analogy. We'll discuss it more as this essay unfolds, but some parallels are listed in the table below: Regular Computer Programs Neural Networks Reverse Engineering Mechanistic Interpretability Program Binary Network Parameters VM / Processor / Interpreter Network Architecture Program State / Memory Layer Representation / Activations Variable / Memory Location Neuron / Feature Direction Taking this analogy seriously can let us explore some of the big picture questions in mechanistic interpretability. Often, questions that feel speculative and slippery for reverse engineering neural networks become clear if you pose the same question for rever
Mechanistic Interpretability, Variables, and the Importance of Interpretable Bases Transformer Circuits Thread Mechanistic Interpretability, Variables, and the Importance of Interpretable Bases An informal note on some intuitions related to Mechanistic Interpretability by Chris Olah. Mechanistic interpretability seeks to reverse engineer neural networks, similar to how one might reverse engineer a compiled binary computer program. After all, neural network parameters are in some sense a binary computer program which runs on one of the exotic virtual machines we call a neural network architectu
Explore this link on the map →related reading
- Mechanistic Interpretability, Variables, and the Importance of Interpretable Basestransformer-circuits.pub
- The Building Blocks of Interpretabilitydistill.pub
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Introduction to Mechanistic Interpretability - by Sarahblog.bluedot.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Interpreting Language Model Parametersgoodfire.ai
- Transformer Circuits Threadtransformer-circuits.pub
- Toy Models of Superpositiontransformer-circuits.pub
- Sparsify: A mechanistic interpretability research agenda — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Interpretability — LessWronglesswrong.com