Dario Amodei — The Urgency of Interpretability
In the decade that I have been working on AI, I’ve watched it grow from a tiny academic field to arguably the most important economic and geopolitical issue in the world. In all that time, perhaps the most important lesson I’ve learned is this: the progress of the underlying technology is inexorable, driven by forces too powerful to stop, but the way in which it happens—the order in which things are built, the applications we choose, and the details of how it is rolled out to society—are eminently possible to change, and it’s possible to have great positive impact by doing so. We can’t stop the bus, but we can steer it. In the past I’ve written about the importance of deploying AI in a way that is positive for the world, and of ensuring that democracies build and wield the technology before autocracies do. Over the last few months, I have become increasingly focused on an additional opportunity for steering the bus: the tantalizing possibility, opened up by some recent advances, th
The Urgency of Interpretability April 2025 In the decade that I have been working on AI, I’ve watched it grow from a tiny academic field to arguably the most important economic and geopolitical issue in the world. In all that time, perhaps the most important lesson I’ve learned is this: the progress of the underlying technology is inexorable, driven by forces too powerful to stop, but the way in which it happens—the order in which things are built, the applications we choose, and the details of how it is rolled out to society—are eminently possible to change, and it’s possible to have great po
Explore this link on the map →saved by
- Phuong Ha Tran Nguyen
- Winnie Xu
- Katherine Huang
- Trang Đoàn
- Justin Wang
- Tasha Pais
- Ratan Kaliani
- HudZah
- Freeman Jiang
- Shivam Siddaiya
- Samuel Lo
- Raffi Hotter
related reading
- On Optimism for Interpretabilitygoodfire.ai
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- The Misguided Quest for Mechanistic AI Interpretability | AI Frontiersai-frontiers.org
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- The Misguided Quest for Mechanistic AI Interpretability | AI Frontiersai-frontiers.org
- Against Almost Every Theory of Impact of Interpretability — AI Alignment Forumalignmentforum.org
- Against Almost Every Theory of Impact of Interpretability — LessWronglesswrong.com
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org