A Mechanistic Interpretability Analysis of Grokking - LessWrong
lesswrong.com · 5,862 words · saved by 1 readers
A significantly updated version of this work is now on Arxiv …
x A Mechanistic Interpretability Analysis of Grokking — LessWrong Best of LessWrong 2022 Interpretability (ML & AI) Lottery Ticket Hypothesis Machine Learning (ML) Grokking (ML) AI Curated 378 A Mechanistic Interpretability Analysis of Grokking by Neel Nanda , Tom Lieberum 15th Aug 2022 AI Alignment Forum Linkpost for colab.research.google.com 43 min read 48 378 Ω 106 A significantly updated version of this work is now on Arxiv and was published as a spotlight paper at ICLR 2023 aka, how the best way to do modular addition is with Discrete Fourier Transforms and trig identities If you don't wa
related reading
- A Mechanistic Interpretability Analysis of Grokking — AI Alignment Forumalignmentforum.org
- Do Machine Learning Models Memorize or Generalize?pair.withgoogle.com
- [2301.05217] Progress measures for grokking via mechanistic interpretabilityarxiv.org
- Towards Understanding Grokking: An Effective Theory of Representation Learningarxiv.org
- Transformer Circuits Threadtransformer-circuits.pub
- Clare Lyle | What's grokking good for?clarelyle.com
- [2301.05217] Progress measures for grokking via mechanistic interpretabilityarxiv.org
- Discovering 108 tricks to accelerate grokkingkindxiaoming.github.io
- In-context Learning and Induction Headstransformer-circuits.pub
- Zipfian grokking | Jasper Gilleyjagilley.github.io
- Mechanistic Interpretability: Circuits, Induction Headsmbrenndoerfer.com
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io