[2301.05217] Progress measures for grokking via mechanistic interpretability
Neural networks often exhibit emergent behavior, where qualitatively new capabilities arise from scaling up the amount of parameters, training data, or training steps. One approach to understanding emergence is to find continuous \textit{progress measures} that underlie the seemingly discontinuous qualitative changes. We argue that progress measures can be found via mechanistic interpretability: reverse-engineering learned behaviors into their individual components. As a case study, we investigate the recently-discovered phenomenon of ``grokking'' exhibited by small transformers trained on modular addition tasks. We fully reverse engineer the algorithm learned by these networks, which uses discrete Fourier transforms and trigonometric identities to convert addition to rotation about a circle. We confirm the algorithm by analyzing the activations and weights and by performing ablations in Fourier space. Based on this understanding, we define progress measures that allow us to study the dynamics of training and split training into three continuous phases: memorization, circuit formation, and cleanup. Our results show that grokking, rather than being a sudden shift, arises from the gradual amplification of structured mechanisms encoded in the weights, followed by the later removal of memorizing components.
Published as a conference paper at ICLR 2023 P ROGRESS MEASURES FOR GROKKING VIA MECHANISTIC INTERPRETABILITY Neel Nanda∗, † Lawrence Chan‡ Tom Lieberum† Jess Smith† Jacob Steinhardt‡ A BSTRACT Neural networks often exhibit emergent behavior, where qualitatively new capa-…
related reading
- [2301.05217] Progress measures for grokking via mechanistic interpretabilityarxiv.org
- A Mechanistic Interpretability Analysis of Grokking — AI Alignment Forumalignmentforum.org
- A Mechanistic Interpretability Analysis of Grokking — LessWronglesswrong.com
- Do Machine Learning Models Memorize or Generalize?pair.withgoogle.com
- Transformer Circuits Threadtransformer-circuits.pub
- On neural scaling and the quanta hypothesisericjmichaud.com
- Zoom In: An Introduction to Circuitsdistill.pub
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- Mechanistic Interpretability: Circuits, Induction Headsmbrenndoerfer.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Towards Automated Circuit Discovery for Mechanistic Interpretabilityarxiv.org
- Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decompositionarxiv.org