[2508.11214] How Causal Abstraction Underpins Computational Explanation
Abstract:Explanations of cognitive behavior often appeal to computations over representations. What does it take for a system to implement a given computation over suitable representational vehicles within that system? We argue that the language of causality -- and specifically the theory of causal abstraction -- provides a fruitful lens on this topic. Drawing on current discussions in deep learning with artificial neural networks, we illustrate how classical themes in the philosophy of computation and cognition resurface in contemporary machine learning. We offer an account of computational implementation grounded in causal abstraction, and examine the role for representation in the resulting picture. We argue that these issues are most profitably explored in connection with generalization and prediction.
HOW CAUSAL ABSTRACTION UNDERPINS COMPUTATIONAL EXPLANATION ATTICUS GEIGER, JACQUELINE HARDING, THOMAS ICARD Abstract. Explanations of cognition often appeal to computations over representations. What does it take for a system to implement a given computation over suitable representa- tional vehicles within that system? We argue that…
saved by
related reading
- How Causal Abstraction Underpins Computational Explanationarxiv.org
- Faithful, Interpretable Model Explanations via Causal Abstraction | SAIL Blogai.stanford.edu
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] — LessWronglesswrong.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Computation in Physical Systems (Stanford Encyclopedia of Philosophy)plato.stanford.edu
- Interpretability Dreamstransformer-circuits.pub
- What If We Had Bigger Brains? Imagining Minds beyond Ours-Stephen Wolfram Writingswritings.stephenwolfram.com
- Training Language Models to Explain Their Own Computationsarxiv.org
- Is causality the missing piece of the AI puzzle?qualcomm.com
- Topicslearnmechinterp.com