flâneur

Mechanistic Interpretability: Circuits, Induction Heads - Interactive | Michael Brenndoerfer

mbrenndoerfer.com · 6,280 words · saved by 1 readers

Reverse-engineer transformer networks into human-understandable algorithms by identifying circuits, induction heads, and mechanistic discoveries.

Mechanistic InterpretabilityLink Copied When a language model predicts that "The Eiffel Tower is located in" should be followed by "Paris," what computation produced that answer? Which weights fired, which attention heads activated, and which neurons collectively encoded the relevant geographic knowledge? Mechanistic interpretability is the research program that asks precisely these questions. Rather than treating a neural network as a black box that produces outputs from inputs, mechanistic interpretability tries to reverse-engineer the model into a human-understandable algorithm. It seeks…

saved by

related reading