flâneur — a map of the web's best reading

[Interim research report] Activation plateaus & sensitive directions in GPT2 — LessWrong

lesswrong.com · 4,355 words · saved by 1 readers

This part-report / part-proposal describes ongoing research, but I'd like to share early results for feedback. I am especially interested in any comment finding mistakes or trivial explanations for these results. I will work on this proposal with a LASR Labs team over the next 3 months. If you are working (or want to work) on something similar I would love to chat! Experiments and write-up by Stefan, with substantial inspiration and advice from Jake (who doesn’t necessarily endorse every sloppy statement I write). Work produced at Apollo Research. TL,DR: Toy models of how neural networks compute new features in superposition seem to imply that neural networks that utilize superposition require some form of error correction to avoid interference spiraling out of control. This means small variations along a feature direction shouldn't affect model outputs, which I can test: I find that both of these predictions hold; the latter when I operationalize "feature" as the difference between tw

x [Interim research report] Activation plateaus & sensitive directions in GPT2 — LessWrong Apollo Research (org) Interpretability (ML & AI) Sparse Autoencoders (SAEs) AI Frontpage 66 [Interim research report] Activation plateaus & sensitive directions in GPT2 by StefanHex , jake_mendel 5th Jul 2024 AI Alignment Forum 7 min read 2 66 Ω 34 This part-report / part-proposal describes ongoing research, but I'd like to share early results for feedback. I am especially interested in any comment finding mistakes or trivial explanations for these results. I will work on this proposal with a LASR Labs t

Explore this link on the map →

related reading