[2601.12703] Towards Spectroscopy: Susceptibility Clusters in Language Models
Abstract:Spectroscopy infers the internal structure of physical systems by measuring their response to perturbations. We apply this principle to neural networks: perturbing the data distribution by upweighting a token $y$ in context $x$, we measure the model's response via susceptibilities $\chi_{xy}$, which are covariances between component-level observables and the perturbation computed over a localized Gibbs posterior via stochastic gradient Langevin dynamics (SGLD). Theoretically, we show that susceptibilities decompose as a sum over modes of the data distribution, explaining why tokens that follow their contexts "for similar reasons" cluster together in susceptibility space. Empirically, we apply this methodology to Pythia-14M, developing a conductance-based clustering algorithm that identifies 510 interpretable clusters ranging from grammatical patterns to code structure to mathematical notation. Comparing to sparse autoencoders, 50% of our clusters match SAE features, validating that both methods recover similar structure.
Towards Spectroscopy: Susceptibility Clusters in Language Models Andrew Gordon= Garrett Baker= George Wang Timaeus Timaeus Timaeus andrew@timaeus.co garrett@timaeus.co george@timaeus.co arXiv:2601.12703v1 [cs.LG] 19 Jan 2026…
saved by
related reading
- Softmax Linear Unitstransformer-circuits.pub
- Neuronpedianeuronpedia.org
- Towards Automated Circuit Discovery for Mechanistic Interpretabilityarxiv.org
- Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decompositionarxiv.org
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Transformer Circuits Threadtransformer-circuits.pub
- Structure and Interpretation of Deep Networkssidn.baulab.info
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub