flâneur — a map of the web's best reading

A short note on interpretability and minds

ericjmichaud.com · 1,916 words · saved by 1 readers

I recently put out a blog post reflecting on a paper from my PhD and its connection to some related work on interpretability and neural scaling laws. Writing that post (intermittently) took me about eighteen months, and it has about 15000 words. In the post below, I've taken a shot at distilling the spirit of that longer post into something much shorter. It is about the assumptions that mechanistic interpretability researchers make and how those assumptions gesture towards a deeper science—a path for machine learning to teach us something new and important about ourselves. It is a bit whimsical, a bit ungrounded, but I nevertheless feel like it is gesturing towards something important and true.

A short note on interpretability and minds A short note on interpretability and minds 2026-04-05 I recently put out a blog post reflecting on a paper from my PhD and its connection to some related work on interpretability and neural scaling laws. Writing that post (intermittently) took me about eighteen months, and it has about 15000 words. In the post below, I've taken a shot at distilling the spirit of that longer post into something much shorter. It is about the assumptions that mechanistic interpretability researchers make and how those assumptions gesture towards a deeper science—a path f

Explore this link on the map →

saved by

related reading