Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] — LessWrong
Summary: This post introduces causal scrubbing, a principled approach for evaluating the quality of mechanistic interpretations. The key idea behind causal scrubbing is to test interpretability hypotheses via behavior-preserving resampling ablations. We apply this method to develop a refined understanding of how a small language model implements induction and how an algorithmic model correctly classifies if a sequence of parentheses is balanced. A question that all mechanistic interpretability work must answer is, “how well does this interpretation explain the phenomenon being studied?”. In the many recent papers in mechanistic interpretability, researchers have generally relied on ad-hoc methods to evaluate the quality of interpretations.[1] This ad hoc nature of existing evaluation methods poses a serious challenge for scaling up mechanistic interpretability. Currently, to evaluate the quality of a particular research result, we need to deeply understand both the interpretation and t
x Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] — LessWrong [Redwood Research] Causal Scrubbing Redwood Research Causal Scrubbing Interpretability (ML & AI) AI Frontpage 208 Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] by LawrenceC , Adrià Garriga-alonso , Nicholas Goldowsky-Dill , ryan_greenblatt , jenny , Ansh Radhakrishnan , Buck , Nate Thomas 3rd Dec 2022 AI Alignment Forum 24 min read 35 208 Ω 103 * Authors sorted alphabetically. Summary: This post introduces causal scrubbing, a principl
Explore this link on the map →saved by
related reading
- Causal scrubbing: Appendix — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Faithful, Interpretable Model Explanations via Causal Abstraction | SAIL Blogai.stanford.edu
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- The Building Blocks of Interpretabilitydistill.pub
- Transformer Circuits Threadtransformer-circuits.pub
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWronglesswrong.com
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org