A short note on interpretability and minds
I recently put out a blog post reflecting on a paper from my PhD and its connection to some related work on interpretability and neural scaling laws. Writing that post (intermittently) took me about eighteen months, and it has about 15000 words. In the post below, I've taken a shot at distilling the spirit of that longer post into something much shorter. It is about the assumptions that mechanistic interpretability researchers make and how those assumptions gesture towards a deeper science—a path for machine learning to teach us something new and important about ourselves. It is a bit whimsical, a bit ungrounded, but I nevertheless feel like it is gesturing towards something important and true.
A short note on interpretability and minds A short note on interpretability and minds 2026-04-05 I recently put out a blog post reflecting on a paper from my PhD and its connection to some related work on interpretability and neural scaling laws. Writing that post (intermittently) took me about eighteen months, and it has about 15000 words. In the post below, I've taken a shot at distilling the spirit of that longer post into something much shorter. It is about the assumptions that mechanistic interpretability researchers make and how those assumptions gesture towards a deeper science—a path f
Explore this link on the map →saved by
related reading
- Interpreting Language Model Parametersgoodfire.ai
- On neural scaling and the quanta hypothesisericjmichaud.com
- Transformer Circuits Threadtransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- The Scaling Hypothesis · Gwern.netgwern.net
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- What is the purpose of interpretability?ericjmichaud.com
- On Optimism for Interpretabilitygoodfire.ai
- The Building Blocks of Interpretabilitydistill.pub
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net