Fact Finding: Simplifying the Circuit (Post 2) — LessWrong
This is the second post in the Google DeepMind mechanistic interpretability team’s investigation into how language models recall facts. This post foc…
x Fact Finding: Simplifying the Circuit (Post 2) — LessWrong Interpretability (ML & AI) Frontpage 27 Fact Finding: Simplifying the Circuit (Post 2) by Senthooran Rajamanoharan , Neel Nanda , János Kramár , Rohin Shah 23rd Dec 2023 AI Alignment Forum 17 min read 3 27 Ω 16 This is the second post in the Google DeepMind mechanistic interpretability team’s investigation into how language models recall facts . This post focuses on distilling down the fact recall circuit and models a more standard mechanistic interpretability investigation. This post gets in the weeds, we recommend starting with pos
related reading
- Fact Finding: Attempting to Reverse-Engineer Factual Recall on the Neuron Level (Post 1) — AI Alignment Forumalignmentforum.org
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Fact Finding: Attempting to Reverse-Engineer Factual Recall on the Neuron Level (Post 1) — LessWronglesswrong.com
- On the Biology of a Large Language Modeltransformer-circuits.pub
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWronglesswrong.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Bridging the Attention Gap: Complete Replacement Models for Complete Circuit Tracinginterp.open-moss.com
- Tiny Mech Interp Projects: Emergent Positional Embeddings of Words - Neel Nandaneelnanda.io