3.1 Importance of Interpretability | Interpretable Machine Learning
Machine learning algorithms usually operate as black boxes and it is unclear how they derived a certain decision. This book is a guide for practitioners to make machine learning decisions interpretable.
This chapter introduces the concepts of interpretability. While it’s difficult to define interpretability mathematically, I like the definition by Biran and Cotton (2017), which was also used by Miller (2019): “Interpretability is the degree to which a human can understand the cause of a decision.” Another good one is by Kim, Khanna, and Koyejo (2016): “a method is interpretable if a user can correctly and efficiently predict the method’s results” The more interpretable a machine learning model, the easier it is for someone to understand why certain decisions or predictions were made. A…
related reading
- 6 – Interpretability – Machine Learning Blog | ML@CMU | Carnegie Mellon Universityblog.ml.cmu.edu
- Against Interpretability: a Critical Examination of the Interpretability Problem in Machine Learninglink.springer.com
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- [1811.10154] Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Insteadarxiv.org
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- On Optimism for Interpretabilitygoodfire.ai
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- In (highly contingent!) defense of interpretability-in-the-loop ML training — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- The Building Blocks of Interpretabilitydistill.pub
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org