[2403.10415] Gradient based Feature Attribution in Explainable AI: A Technical Review
The surge in black-box AI models has prompted the need to explain the internal mechanism and justify their reliability, especially in high-stakes applications, such as healthcare and autonomous driving. Due to the lack…
Gradient based Feature Attribution in Explainable AI: A Technical Review Yongjie Wang Nanyang technological university yongjie002@e.ntu.edu.sg , Tong Zhang Nanyang technological university tong.zhang@ntu.edu.sg , Xu Guo Nanyang technological university xu.guo@ntu.edu.sg and Zhiqi Shen Nanyang technological university zqshen@ntu.edu.sg (2024) Abstract. The surge in black-box AI models has prompted the need to explain the internal mechanism and justify their reliability, especially in high-stakes applications, such as healthcare and autonomous driving. Due to the lack of a rigorous definition of
related reading
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Faithful, Interpretable Model Explanations via Causal Abstraction | SAIL Blogai.stanford.edu
- The Building Blocks of Interpretabilitydistill.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- Feature Visualizationdistill.pub
- Topicslearnmechinterp.com
- Attribution-based parameter decomposition — LessWronglesswrong.com
- ADAG: Automatically Describing Attribution Graphsarxiv.org