Mechanistic Transparency for Machine Learning
[EDIT (added Jan 2023): it’s come to my attention that this post was likely influenced by conversations I had with Chris Olah related to the distinction between standard interpretability and the type called “mechanistic”, as well as early experiments he had which became the ‘circuits’ sequence of papers - my sincere apologies for not making this clearer earlier.]
Mechanistic Transparency for Machine Learning Mechanistic Transparency for Machine Learning Jul 10, 2018 Cross-posted to the AI Alignment Forum . [EDIT (added Jan 2023): it’s come to my attention that this post was likely influenced by conversations I had with Chris Olah related to the distinction between standard interpretability and the type called “mechanistic”, as well as early experiments he had which became the ‘circuits’ sequence of papers - my sincere apologies for not making this clearer earlier.] Lately I’ve been trying to come up with a thread of AI alignment research that (a) I can
Explore this link on the map →related reading
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- On Optimism for Interpretabilitygoodfire.ai
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com
- Introduction to Mechanistic Interpretability - by Sarahblog.bluedot.org
- Interpretability — LessWronglesswrong.com
- Sparsify: A mechanistic interpretability research agenda — AI Alignment Forumalignmentforum.org