SAE feature geometry is outside the superposition hypothesis — LessWrong
Summary: Superposition-based interpretations of neural network activation spaces are incomplete. The specific locations of feature vectors contain cr…
x SAE feature geometry is outside the superposition hypothesis — LessWrong Interpretability (ML & AI) Apollo Research (org) AI Frontpage 229 SAE feature geometry is outside the superposition hypothesis by jake_mendel 24th Jun 2024 AI Alignment Forum 14 min read 18 229 Ω 101 Written at Apollo Research Summary: Superposition-based interpretations of neural network activation spaces are incomplete. The specific locations of feature vectors contain crucial structural information beyond superposition, as seen in circular arrangements of day-of-the-week features and in the rich structures of feature
Explore this link on the map →related reading
- SAE feature geometry is outside the superposition hypothesis — AI Alignment Forumalignmentforum.org
- Toy Models of Superpositiontransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Activation space interpretability may be doomed — LessWronglesswrong.com
- Interpretability Dreamstransformer-circuits.pub
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- pdfopenreview.net
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Lucius Bushnaq's Shortform — LessWronglesswrong.com
- What Would Non-Linear Features Actually Look Like? — Liv Gortonlivgorton.com
- The World Inside Neural Networksgoodfire.ai
- how neural networks think at scalemarmik.xyz