Superposition, Memorization, and Double Descent
In a recent paper , we found that simple neural networks trained on toy tasks often exhibit a phenomenon called superposition , where they represent more features than they have neurons. Our investigation was limited to the infinite-data, underfitting regime. But there's reason to believe that understanding overfitting might be important if we want to succeed at mechanistic interpretability, and that superposition might be a central part of the story. Why should mechanistic interpretability care about overfitting? Despite overfitting being a central problem in machine learning, we have little mechanistic understanding of what exactly is going on when deep learning models overfit or memorize examples. Additionally, previous work has hinted that there may be an important link between overfitting and learning interpretable features . So understanding overfitting is important, but why should it be relevant to superposition? Consider the case of a language model which verbatim memorizes tex
Superposition, Memorization, and Double Descent Transformer Circuits Thread Superposition, Memorization, and Double Descent Authors Tom Henighan ∗ , Shan Carter ∗ , Tristan Hume ∗ , Nelson Elhage ∗ , Robert Lasenby, Stanislav Fort, Nicholas Schiefer, Christopher Olah ‡ Affiliation Anthropic Published January 5, 2023 * Core Research Contributor; ‡ Correspondence to colah@anthropic.com ; Author contributions statement below . In a recent paper , we found that simple neural networks trained on toy tasks often exhibit a phenomenon called superposition , where they represent more features than they
related reading
- Toy Models of Superpositiontransformer-circuits.pub
- Interpretability Dreamstransformer-circuits.pub
- Do Machine Learning Models Memorize or Generalize?pair.withgoogle.com
- Zoom In: An Introduction to Circuitsdistill.pub
- [2303.14151] Double Descent Demystified: Identifying, Interpreting & Ablating the Sources of a Deep Learning Puzzlearxiv.org
- Understanding “Deep Double Descent” — LessWronglesswrong.com
- What Would Non-Linear Features Actually Look Like? — Liv Gortonlivgorton.com
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- On neural scaling and the quanta hypothesisericjmichaud.com
- Zipfian grokking | Jasper Gilleyjagilley.github.io
- Clare Lyle | What's grokking good for?clarelyle.com
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasetsmathai-iclr.github.io