✳flâneur — a map of the web's best reading
Circuit Tracing: Revealing Computational Graphs in Language Models
transformer-circuits.pub · 29,625 words · saved by 18 readers
We describe an approach to tracing the “step-by-step” computation involved when a model responds to a single prompt.
Circuit Tracing: Revealing Computational Graphs in Language Models × Transformer Circuits Thread Circuit Tracing: Revealing Computational Graphs in Language Models Circuit Tracing: Revealing Computational Graphs in Language Models We introduce a method to uncover mechanisms underlying behaviors of language models. We produce graph descriptions of the model’s computation on prompts of interest by tracing individual computational steps in a “replacement model”. This replacement model substitutes a more interpretable component (here, a “cross-layer transcoder”) for parts of the underlying model (
Explore this link on the map →saved by
- Karan MJ
- Tasha Pais
- Samuel Lo
- Jennifer Zhao
- Matthew Wang
- Asher P
- Vyom Pathak
- Thu Than
- Lydia Nottingham
- Logan Graves
- Eric Huang
- Cheikh Fiteni
related reading
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWronglesswrong.com
- Language Model Circuits Are Sparse in the Neuron Basis | Transluce AItransluce.org
- Tracing Attention Computation Through Feature Interactionstransformer-circuits.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org