flâneur — a map of the web's best reading

Language Model Circuits Are Sparse in the Neuron Basis | Transluce AI

transluce.org · 18,651 words · saved by 1 readers

We introduce a new approach to finding sparse circuits in the neuron basis, without relying on learned features.

Language Model Circuits Are Sparse in the Neuron Basis Aryaman Arora * , Zhengxuan Wu * , Jacob Steinhardt , Sarah Schwettmann * Equal contribution. Correspondence to: aryaman@transluce.org, zen@transluce.org. Transluce | Published: November 20, 2025 Many interpretability methods rely on learned feature bases—such as sparse autoencoders or cross-layer transcoders—based on the belief that neurons do not cleanly decompose model computation. We revisit this assumption and show that, with a better choice of neuron basis (MLP activations) and a stronger attribution method (RelP), raw neurons can pr

Explore this link on the map →

related reading