[2608.27540] Towards a mathematical theory of superposition
Abstract:We develop a mathematical theory of superposition in neural networks using tools from frame theory and compressed sensing. In our model, a sparse binary vector \(x\) of active features is encoded through an overcomplete dictionary \(W\), and feature recovery is performed by applying \(\operatorname{ReLU}(W^\top W x+b)\) with an appropriate bias vector \(b\). We prove several recovery theorems for this model. In the random-support setting, we establish high-probability support recovery for nearly tight, low-coherence dictionaries, with guarantees when the expected sparsity is up to order \(d/\log n\). In the worst-case support setting, we give a sharp and computable criterion for which sparsity levels permit support recovery. We apply this criterion to Gaussian random matrices and equiangular tight frames. For real equiangular tight frames with \(n>d+1\), we determine the exact recovery threshold in terms of the coherence. The proof of this result for real equiangular tight frames relies on a novel characterization---which should be of independent interest to frame theorists---of the distribution of signs in the Gram matrix.
Towards a mathematical theory of superposition Michael I. Ivanitskiy∗ John Jasper† Emily J. King‡§ Dustin G. Mixon¶‖ Abstract We develop a mathematical theory of superposition in neural networks using tools from frame theory and compressed sensing. In our model, a sparse binary vector x of active features is…
saved by
related reading
- [2605.01192] Linear-Readout Floors and Threshold Recovery in Computation in Superpositionarxiv.org
- Toy Models of Superpositiontransformer-circuits.pub
- k-Sparse Autoencodersarxiv.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- What Would Non-Linear Features Actually Look Like? — Liv Gortonlivgorton.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- [Interim research report] Taking features out of superposition with sparse autoencoders — LessWronglesswrong.com
- Circuits in Superposition 2: Now with Less Wrong Math — LessWronglesswrong.com
- Interpretability Dreamstransformer-circuits.pub
- [2410.12101] The Persian Rug: solving toy models of superposition using large-scale symmetriesarxiv.org
- Computational Superposition in a Toy Model of the U-AND Problem — LessWronglesswrong.com
- Distributed Representations: Composition & Superpositiontransformer-circuits.pub