A short note on interpretability and minds
I recently put out a blog post reflecting on a paper from my PhD and its connection to some related work on interpretability and neural scaling laws. Writing that post (intermittently) took me about eighteen months, and it has about 15000 words. In the post below, I've taken a shot at distilling the spirit of that longer post into something much shorter. It is about the assumptions that mechanistic interpretability researchers make and how those assumptions gesture towards a deeper science—a path for machine learning to teach us something new and important about ourselves. It is a bit whimsical, a bit ungrounded, but I nevertheless feel like it is gesturing towards something important and true.
A short note on interpretability and minds A short note on interpretability and minds 2026-04-05 I recently put out a blog post reflecting on a paper from my PhD and its connection to some related work on interpretability and neural scaling laws. Writing that post (intermittently) took me about eighteen months, and it has about 15000 words. In the post below, I've taken a shot at distilling the spirit of that longer post into something much shorter. It is about the assumptions that mechanistic interpretability researchers make and how those assumptions gesture towards a deeper science—a path f
saved by
related reading
- On neural scaling and the quanta hypothesisericjmichaud.com
- Interpreting Language Model Parametersgoodfire.ai
- Transformer Circuits Threadtransformer-circuits.pub
- What is the purpose of interpretability?ericjmichaud.com
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- The Scaling Hypothesis · Gwern.netgwern.net
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- On Optimism for Interpretabilitygoodfire.ai
- The Building Blocks of Interpretabilitydistill.pub