✳flâneur — a map of the web's best reading
A Mechanistic Interpretability Analysis of Grokking - AI Alignment Forum
alignmentforum.org · 7,878 words · saved by 1 readers
Neel Nanda reverse engineers neural networks that have "grokked" modular addition, showing that they operate using Discrete Fourier Transforms and tr…
x A Mechanistic Interpretability Analysis of Grokking — AI Alignment Forum Best of LessWrong 2022 Interpretability (ML & AI) Lottery Ticket Hypothesis Machine Learning (ML) Grokking (ML) AI Curated 106 A Mechanistic Interpretability Analysis of Grokking by Neel Nanda , Tom Lieberum 15th Aug 2022 Linkpost for colab.research.google.com 43 min read 48 106 A significantly updated version of this work is now on Arxiv and was published as a spotlight paper at ICLR 2023 aka, how the best way to do modular addition is with Discrete Fourier Transforms and trig identities If you don't want to commit to
Explore this link on the map →related reading
- A Mechanistic Interpretability Analysis of Grokking — LessWronglesswrong.com
- Do Machine Learning Models Memorize or Generalize?pair.withgoogle.com
- [2301.05217] Progress measures for grokking via mechanistic interpretabilityarxiv.org
- [2301.05217] Progress measures for grokking via mechanistic interpretabilityarxiv.org
- Transformer Circuits Threadtransformer-circuits.pub
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Clare Lyle | What's grokking good for?clarelyle.com
- [2602.16849] On the Mechanism and Dynamics of Modular Addition: Fourier Features, Lottery Ticket, and Grokkingarxiv.org
- On neural scaling and the quanta hypothesisericjmichaud.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Zipfian grokking | Jasper Gilleyjagilley.github.io