Mechanistic Interpretability: A Challenge Common to Both Artificial and Biological Intelligence - Kempner Institute
In neuroscience, the past decade has witnessed major advances in our ability to record activity from the brain at both larger and finer scales. And yet, a mechanistic theory linking […]
In neuroscience, the past decade has witnessed major advances in our ability to record activity from the brain at both larger and finer scales. And yet, a mechanistic theory linking computation to intelligent behavior remains elusive. By “mechanistic theory,” we mean explanations, in terms of human-understandable concepts, of the activity of neurons and how they relate to an agent’s behavior. Mechanistic theories seem to have become even more elusive as experimental research has been shifting from highly-structured research, where experimentalists can exercise explicit control over…
saved by
related reading
- Computational and systems neuroscience: The next 20 yearspmc.ncbi.nlm.nih.gov
- Trading places: What happens when neuroscience turns into machine learning, and machine learning turns into neuroscience?thetransmitter.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Introduction to Mechanistic Interpretability - by Sarahblog.bluedot.org
- On Optimism for Interpretabilitygoodfire.ai
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io
- Sparsify: A mechanistic interpretability research agenda — AI Alignment Forumalignmentforum.org
- The Building Blocks of Interpretabilitydistill.pub
- Perfectly Normalperfectlynormal.co.uk