[Interim research report] Activation plateaus & sensitive directions in GPT2 — LessWrong
This part-report / part-proposal describes ongoing research, but I'd like to share early results for feedback. I am especially interested in any comment finding mistakes or trivial explanations for these results. I will work on this proposal with a LASR Labs team over the next 3 months. If you are working (or want to work) on something similar I would love to chat! Experiments and write-up by Stefan, with substantial inspiration and advice from Jake (who doesn’t necessarily endorse every sloppy statement I write). Work produced at Apollo Research. TL,DR: Toy models of how neural networks compute new features in superposition seem to imply that neural networks that utilize superposition require some form of error correction to avoid interference spiraling out of control. This means small variations along a feature direction shouldn't affect model outputs, which I can test: I find that both of these predictions hold; the latter when I operationalize "feature" as the difference between tw
x [Interim research report] Activation plateaus & sensitive directions in GPT2 — LessWrong Apollo Research (org) Interpretability (ML & AI) Sparse Autoencoders (SAEs) AI Frontpage 66 [Interim research report] Activation plateaus & sensitive directions in GPT2 by StefanHex , jake_mendel 5th Jul 2024 AI Alignment Forum 7 min read 2 66 Ω 34 This part-report / part-proposal describes ongoing research, but I'd like to share early results for feedback. I am especially interested in any comment finding mistakes or trivial explanations for these results. I will work on this proposal with a LASR Labs t
Explore this link on the map →related reading
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Toy Models of Superpositiontransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- Steering GPT-2-XL by adding an activation vector — AI Alignment Forumalignmentforum.org
- SAE feature geometry is outside the superposition hypothesis — AI Alignment Forumalignmentforum.org
- Neuronpedianeuronpedia.org
- Lucius Bushnaq's Shortform — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- SAE feature geometry is outside the superposition hypothesis — LessWronglesswrong.com
- Activation space interpretability may be doomed — LessWronglesswrong.com
- Softmax Linear Unitstransformer-circuits.pub