Causal scrubbing: Appendix — LessWrong
lesswrong.com · 7,613 words · saved by 1 readers
* Authors sorted alphabetically. • An appendix to this post. • 1 More on Hypotheses …
x Causal scrubbing: Appendix — LessWrong [Redwood Research] Causal Scrubbing Causal Scrubbing Interpretability (ML & AI) Redwood Research AI Frontpage 18 Causal scrubbing: Appendix by LawrenceC , Adrià Garriga-alonso , Nicholas Goldowsky-Dill , ryan_greenblatt , jenny , Ansh Radhakrishnan , Buck , Nate Thomas 3rd Dec 2022 AI Alignment Forum 24 min read 4 18 Ω 11 * Authors sorted alphabetically. An appendix to this post . 1 More on Hypotheses 1.1 Example behaviors As mentioned above, our method allows us to explain quantitatively measured model behavior operationalized as the expectation of a f
saved by
related reading
- Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] — LessWronglesswrong.com
- Towards Automated Circuit Discovery for Mechanistic Interpretabilityarxiv.org
- Faithful, Interpretable Model Explanations via Causal Abstraction | SAIL Blogai.stanford.edu
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Topicslearnmechinterp.com
- Weight-Sparse Circuits May Be Interpretable Yet Unfaithful — LessWronglesswrong.com
- Interpretability Dreamstransformer-circuits.pub
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- [2508.11214] How Causal Abstraction Underpins Computational Explanationarxiv.org
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com