Uncovering Neural Geometry in Vision Models With Block-Sparse Featurizers
We introduce Block-Sparse Featurizers (BSF), a family of methods to decompose a model’s activations into multidimensional subspaces rather than single directions. Applied to vision models, we find that BSFs find interpretable, multidimensional features which offer a more parsimonious explanation of model internals; that those features enable fine-grained steering; and that most concepts in the models are multidimensional.
Uncovering Neural Geometry in Vision Models With Block-Sparse Featurizers Research ← The Neural Geometry Series Uncovering Neural Geometry in Vision Models With Block-Sparse Featurizers We introduce Block-Sparse Featurizers (BSF), a family of methods to decompose a model's activations into multidimensional subspaces rather than single directions. Applied to vision models, we find that BSFs find interpretable, multidimensional features which offer a more parsimonious explanation of model internals; that those features enable fine-grained steering; and that most concepts in the models are multid
Explore this link on the map →related reading
- Toy Models of Superpositiontransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- The World Inside Neural Networksgoodfire.ai
- Feature Visualizationdistill.pub
- The Building Blocks of Interpretabilitydistill.pub
- Activation space interpretability may be doomed — LessWronglesswrong.com
- Feature Visualizationdistill.pub
- how neural networks think at scalemarmik.xyz
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- What Would Non-Linear Features Actually Look Like? — Liv Gortonlivgorton.com
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io