A Mechanistic Interpretability Analysis of Grokking - AI Alignment Forum
alignmentforum.org · 7,878 words · saved by 1 readers
Neel Nanda reverse engineers neural networks that have "grokked" modular addition, showing that they operate using Discrete Fourier Transforms and tr…
x A Mechanistic Interpretability Analysis of Grokking — AI Alignment Forum Best of LessWrong 2022 Interpretability (ML & AI) Lottery Ticket Hypothesis Machine Learning (ML) Grokking (ML) AI Curated 106 A Mechanistic Interpretability Analysis of Grokking by Neel Nanda , Tom Lieberum 15th Aug 2022 Linkpost for colab.research.google.com 43 min read 48 106 A significantly updated version of this work is now on Arxiv and was published as a spotlight paper at ICLR 2023 aka, how the best way to do modular addition is with Discrete Fourier Transforms and trig identities If you don't want to commit to
related reading
- A Mechanistic Interpretability Analysis of Grokking — LessWronglesswrong.com
- Do Machine Learning Models Memorize or Generalize?pair.withgoogle.com
- [2301.05217] Progress measures for grokking via mechanistic interpretabilityarxiv.org
- Transformer Circuits Threadtransformer-circuits.pub
- [2301.05217] Progress measures for grokking via mechanistic interpretabilityarxiv.org
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Mechanistic Interpretability: Circuits, Induction Headsmbrenndoerfer.com
- [2602.16849] On the Mechanism and Dynamics of Modular Addition: Fourier Features, Lottery Ticket, and Grokkingarxiv.org
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Clare Lyle | What's grokking good for?clarelyle.com
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- Towards Understanding Grokking: An Effective Theory of Representation Learningarxiv.org