Garcon
This document describes one of the core pieces of infrastructure we've built to enable interpretability research at Anthropic: a tool we call Garçon. Among other research projects, at Anthropic we work on interpretability -- trying to understand what is going on inside of language models as they work. We take a number of approaches, but we spend a lot of time on “Circuits”-style mechanistic interpretability, trying to dig deeply into the actual mechanics of the computation performed by a model. This kind of work involves “poking at” models in a broad and flexible way. We want to examine individual activations inside arbitrary layers of the model; to perform experiments where we modify or ablate individual components; and otherwise to access and work with the “guts” of the model, and not just its inputs and outputs. We also desire to scale this work up to the largest models that Anthropic works with. Modern Transformer models can be very large indeed, and large models can be so large th
Garcon Garcon Authors Nelson Elhage ∗ , Neel Nanda , Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann , Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds , Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, Chris Olah Affiliation Anthropic Published Dec 22, 2021 * Core Contributor You can also watch a video covering similar content to this piece. This document describes one of the core pieces of infrastructure we've built to enable int
Explore this link on the map →related reading
- Transformer Circuits Threadtransformer-circuits.pub
- Interpretability Infrastructure at Frontier Scale: Harvesting Activations from a Trillion-Parameter Modelgoodfire.ai
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- The Building Blocks of Interpretabilitydistill.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Prism: mapping interpretable concepts and features in a latent space of language | thesephist.comthesephist.com
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- Neuronpedianeuronpedia.org
- TransformerLens Documentationtransformerlensorg.github.io