6 – Interpretability – Machine Learning Blog | ML@CMU | Carnegie Mellon University
The objectives machine learning models optimize for do not always reflect the actual desiderata of the task at hand. Interpretability in models allows us to evaluate their decisions and obtain information that the objective alone cannot confer. Interpretability takes many forms and can be difficult
6 – Interpretability – Machine Learning Blog | ML@CMU | Carnegie Mellon University --> Input your search keywords and press Enter. Categories: Research Educational Categories: Research Educational Educational machine learning 6 – Interpretability Authors Audrey Huang , Jeffrey Li and Naveen Shankar by Audrey Huang --> Affiliations Published August 31, 2020 DOI Figure 1: Interpretability for machine learning models bridges the concrete objectives models optimize for and the real-world (and less easy to define) desiderata that ML applications aim to achieve. Introduction The ob
Explore this link on the map →related reading
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- The Building Blocks of Interpretabilitydistill.pub
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- On Optimism for Interpretabilitygoodfire.ai
- Transformer Circuits Threadtransformer-circuits.pub
- Faithful, Interpretable Model Explanations via Causal Abstraction | SAIL Blogai.stanford.edu
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- Shapley value: from cooperative game to explainable artificial intelligence | Autonomous Intelligent Systems | Springer Nature Linklink.springer.com