Lucius Bushnaq's Shortform — LessWrong
Potentially up to some small ϵ noise. For a nice operationalisation, see definition 2 on page 3 of this paper. It's a vector because we've already assumed that features are all scalar. If a feature was two-dimensional instead, this would be a projection into an associated two-dimensional subspace. I'm using the term basis loosely here, this also includes sparse overcomplete 'bases' like those in SAEs. The more accurate term would probably be 'dictionary', or 'frame'. Or if the computation isn't layer aligned, the activations along some other causal cut through the network can be written as a sum of all the features represented on that cut. Less, if you want to be able to perform computation in superposition. EDIT: I am now more awake. I still think this is right. Relative to the animal features at least. They could still be sparse relative to the rest of the network if this 50-dimensional animal subspace is rarely used. Like, say, politicians. Or natsec people. Outer alignment in the
x Lucius Bushnaq's Shortform — LessWrong Lucius Bushnaq's Shortform by Lucius Bushnaq 6th Jul 2024 1 min read 105 8 This is a special post for quick takes by Lucius Bushnaq . Only they can create top-level comments. Comments here also appear on the Quick Takes page and All Posts page . Lucius Bushnaq's Shortform 90 Lucius Bushnaq 10 leogao 4 Neel Nanda 4 Alexander Gietelink Oldenziel 8 ryan_greenblatt 2 Lucius Bushnaq 2 Sodium 2 Lucius Bushnaq 1 Sodium 1 keith_wynroe 3 Lucius Bushnaq 2 Neel Nanda 87 Lucius Bushnaq 4 Logan Riggs 2 Lucius Bushnaq 4 chanind 6 Lucius Bushnaq 4 J Bostock 2 Lucius B
Explore this link on the map →related reading
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Toy Models of Superpositiontransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- Activation space interpretability may be doomed — LessWronglesswrong.com
- Interpretability Dreamstransformer-circuits.pub
- Circuits Updates - January 2024transformer-circuits.pub