✳flâneur — a map of the web's best reading
[2404.15255] How to use and interpret activation patching
arxiv.org · saved by 1 readers
Abstract:Activation patching is a popular mechanistic interpretability technique, but has many subtleties regarding how it is applied and how one may interpret the results. We provide a summary of advice and best practices, based on our experience using this technique in practice. We include an overview of the different ways to apply activation patching and a discussion on how to interpret the results. We focus on what evidence patching experiments provide about circuits, and on the choice of metric and associated pitfalls.
Explore this link on the map →related reading
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- The Building Blocks of Interpretabilitydistill.pub
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumneelnanda.io
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- Leela Interpretabilityleela-interp.github.io
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com