flâneur — a map of the web's best reading

Garcon

transformer-circuits.pub · 2,058 words · saved by 1 readers

This document describes one of the core pieces of infrastructure we've built to enable interpretability research at Anthropic: a tool we call Garçon. Among other research projects, at Anthropic we work on interpretability -- trying to understand what is going on inside of language models as they work. We take a number of approaches, but we spend a lot of time on “Circuits”-style mechanistic interpretability, trying to dig deeply into the actual mechanics of the computation performed by a model. This kind of work involves “poking at” models in a broad and flexible way. We want to examine individual activations inside arbitrary layers of the model; to perform experiments where we modify or ablate individual components; and otherwise to access and work with the “guts” of the model, and not just its inputs and outputs. We also desire to scale this work up to the largest models that Anthropic works with. Modern Transformer models can be very large indeed, and large models can be so large th

Garcon Garcon Authors Nelson Elhage ∗ , Neel Nanda , Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann , Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds , Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, Chris Olah Affiliation Anthropic Published Dec 22, 2021 * Core Contributor You can also watch a video covering similar content to this piece. This document describes one of the core pieces of infrastructure we've built to enable int

Explore this link on the map →

related reading