Superposition, Memorization, and Double Descent
In a recent paper , we found that simple neural networks trained on toy tasks often exhibit a phenomenon called superposition , where they represent more features than they have neurons. Our investigation was limited to the infinite-data, underfitting regime. But there's reason to believe that understanding overfitting might be important if we want to succeed at mechanistic interpretability, and that superposition might be a central part of the story. Why should mechanistic interpretability care about overfitting? Despite overfitting being a central problem in machine learning, we have little mechanistic understanding of what exactly is going on when deep learning models overfit or memorize examples. Additionally, previous work has hinted that there may be an important link between overfitting and learning interpretable features . So understanding overfitting is important, but why should it be relevant to superposition? Consider the case of a language model which verbatim memorizes tex
Superposition, Memorization, and Double Descent Transformer Circuits Thread Superposition, Memorization, and Double Descent Authors Tom Henighan ∗ , Shan Carter ∗ , Tristan Hume ∗ , Nelson Elhage ∗ , Robert Lasenby, Stanislav Fort, Nicholas Schiefer, Christopher Olah ‡ Affiliation Anthropic Published January 5, 2023 * Core Research Contributor; ‡ Correspondence to colah@anthropic.com ; Author contributions statement below . In a recent paper , we found that simple neural networks trained on toy tasks often exhibit a phenomenon called superposition , where they represent more features than they
Explore this link on the map →related reading
- Toy Models of Superpositiontransformer-circuits.pub
- Interpretability Dreamstransformer-circuits.pub
- Zoom In: An Introduction to Circuitsdistill.pub
- What Would Non-Linear Features Actually Look Like? — Liv Gortonlivgorton.com
- Understanding “Deep Double Descent” — LessWronglesswrong.com
- On neural scaling and the quanta hypothesisericjmichaud.com
- Zipfian grokking | Jasper Gilleyjagilley.github.io
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- Clare Lyle | What's grokking good for?clarelyle.com
- Circuits Updates — May 2023transformer-circuits.pub
- Human-like Neural Nets by Catapulting · Gwern.netgwern.net
- Do Machine Learning Models Memorize or Generalize?pair.withgoogle.com