Run: A/1 | Dictionary Learning
As described in the paper, the visualization below decomposes neural network activations into features using a sparse autoencoder. 31, 95, 122, 131, 135, 191, 223, 225, 227, 240, 260, 309, 325, 368, 414, 433, 445, 446, 448, 459, 574, 581, 588, 626, 695, 892, 940, 1011, 1048, 1072, 1109, 1112, 1221, 1247, 1252, 1289, 1290, 1298, 1302, 1331, 1336, 1370, 1402, 1407, 1427, 1468, 1492, 1504, 1523, 1530, 1552, 1594, 1675, 1701, 1705, 1710, 1744, 1748, 1756, 1765, 1767, 1816, 1831, 1839, 1854, 1860, 1869, 1883, 1884, 1888, 1897, 1966, 1976, 1989, 2029, 2071, 2080, 2086, 2093, 2120, 2218, 2223, 2238, 2246, 2247, 2297, 2333, 2355, 2392, 2408, 2437, 2441, 2447, 2470, 2475, 2511, 2541, 2544, 2551, 2564, 2575, 2583, 2625, 2647, 2670, 2682, 2685, 2692, 2716, 2717, 2726, 2789, 2833, 2869, 2879, 2890, 2938, 2980, 2987, 2992, 3025, 3037, 3051, 3079, 3086, 3113, 3127, 3153, 3171, 3186, 3187, 3202, 3254, 3285, 3296, 3312, 3360, 3377, 3409, 3431, 3455, 3480, 3531, 3538, 3541, 3544, 3626, 3645, 3654, 3661
As described in the paper, the visualization below decomposes neural network activations into features using a sparse autoencoder. 31, 95, 122, 131, 135, 191, 223, 225, 227, 240, 260, 309, 325, 368, 414, 433, 445, 446, 448, 459, 574, 581, 588, 626, 695, 892, 940, 1011, 1048, 1072, 1109, 1112, 1221, 1247, 1252, 1289, 1290, 1298, 1302, 1331, 1336, 1370, 1402, 1407, 1427, 1468, 1492, 1504, 1523, 1530, 1552, 1594, 1675, 1701, 1705, 1710, 1744, 1748, 1756, 1765, 1767, 1816, 1831, 1839, 1854, 1860, 1869, 1883, 1884, 1888, 1897, 1966, 1976, 1989, 2029, 2071, 2080, 2086, 2093, 2120, 2218, 2223, 2238,
Explore this link on the map →related reading
- Run: A/29 | Dictionary Learningtransformer-circuits.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Toy Models of Superpositiontransformer-circuits.pub
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- Feature Visualizationdistill.pub
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Zoom In: An Introduction to Circuitsdistill.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Distill — Latest articles about machine learningdistill.pub
- Activation space interpretability may be doomed — LessWronglesswrong.com
- The Building Blocks of Interpretabilitydistill.pub
- Feature Visualizationdistill.pub