Run: A/29 | Dictionary Learning
As described in the paper, the visualization below decomposes neural network activations into features using a sparse autoencoder. 1, 2, 4, 6, 7, 9, 11, 13, 14, 15, 16, 17, 19, 20, 22, 23, 24, 25, 27, 29, 30, 32, 36, 38, 39, 41, 42, 46, 47, 50, 51, 52, 53, 55, 57, 61, 62, 63, 64, 65, 67, 70, 71, 72, 75, 76, 77, 78, 79, 80, 81, 83, 84, 88, 91, 92, 94, 95, 96, 97, 101, 102, 103, 104, 105, 106, 107, 109, 110, 114, 115, 119, 120, 126, 127, 128, 130, 131, 133, 134, 141, 143, 146, 147, 148, 149, 150, 151, 153, 157, 158, 159, 168, 170, 171, 172, 174, 175, 179, 181, 183, 187, 189, 191, 194, 197, 198, 199, 200, 202, 205, 206, 208, 209, 214, 215, 219, 220, 223, 224, 227, 228, 234, 242, 243, 244, 248, 250, 260, 262, 266, 267, 269, 270, 271, 276, 278, 279, 280, 281, 283, 287, 294, 295, 299, 300, 301, 302, 303, 304, 305, 308, 309, 311, 312, 313, 316, 319, 330, 331, 334, 339, 341, 344, 345, 348, 354, 355, 359, 360, 362, 363, 366, 370, 371, 373, 374, 375, 377, 378, 380, 381, 383, 384, 388, 389, 390,
As described in the paper, the visualization below decomposes neural network activations into features using a sparse autoencoder. 1, 2, 4, 6, 7, 9, 11, 13, 14, 15, 16, 17, 19, 20, 22, 23, 24, 25, 27, 29, 30, 32, 36, 38, 39, 41, 42, 46, 47, 50, 51, 52, 53, 55, 57, 61, 62, 63, 64, 65, 67, 70, 71, 72, 75, 76, 77, 78, 79, 80, 81, 83, 84, 88, 91, 92, 94, 95, 96, 97, 101, 102, 103, 104, 105, 106, 107, 109, 110, 114, 115, 119, 120, 126, 127, 128, 130, 131, 133, 134, 141, 143, 146, 147, 148, 149, 150, 151, 153, 157, 158, 159, 168, 170, 171, 172, 174, 175, 179, 181, 183, 187, 189, 191, 194, 197, 198,
related reading
- Run: A/1 | Dictionary Learningtransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Toy Models of Superpositiontransformer-circuits.pub
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- Feature Visualizationdistill.pub
- Uncovering Neural Geometry in Vision Models With Block-Sparse Featurizersgoodfire.ai
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- sparseAutoencoder.pdfweb.stanford.edu
- Activation space interpretability may be doomed — LessWronglesswrong.com
- [2309.08600] Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org