flâneur — a map of the web's best reading

Superposition, Memorization, and Double Descent

transformer-circuits.pub · 7,217 words · saved by 1 readers

In a recent paper , we found that simple neural networks trained on toy tasks often exhibit a phenomenon called superposition , where they represent more features than they have neurons. Our investigation was limited to the infinite-data, underfitting regime. But there's reason to believe that understanding overfitting might be important if we want to succeed at mechanistic interpretability, and that superposition might be a central part of the story. Why should mechanistic interpretability care about overfitting? Despite overfitting being a central problem in machine learning, we have little mechanistic understanding of what exactly is going on when deep learning models overfit or memorize examples. Additionally, previous work has hinted that there may be an important link between overfitting and learning interpretable features . So understanding overfitting is important, but why should it be relevant to superposition? Consider the case of a language model which verbatim memorizes tex

Superposition, Memorization, and Double Descent Transformer Circuits Thread Superposition, Memorization, and Double Descent Authors Tom Henighan ∗ , Shan Carter ∗ , Tristan Hume ∗ , Nelson Elhage ∗ , Robert Lasenby, Stanislav Fort, Nicholas Schiefer, Christopher Olah ‡ Affiliation Anthropic Published January 5, 2023 * Core Research Contributor; ‡ Correspondence to colah@anthropic.com ; Author contributions statement below . In a recent paper , we found that simple neural networks trained on toy tasks often exhibit a phenomenon called superposition , where they represent more features than they

Explore this link on the map →

related reading