Latest research - Goodfire
goodfire.com · 570 words · saved by 1 readers
Our latest interpretability research to understand and intentionally design advanced AI systems.
Filter By · August 20, 2026 Bigelow et al. · August 20, 2026 Fel et al. · July 7, 2026 Fel et al. · July 7, 2026 Bigelow et al. · June 23, 2026 Bigelow et al. · June 23, 2026 Bergen et al. · June 11, 2026 Bergen et al. · June 11, 2026 Santiago Aranguri · June 4, 2026 Santiago Aranguri · June 4, 2026 Huang et al. · June 1, 2026 Huang et al. · June 1, 2026 Bhalla et al. · May 21, 2026 Bhalla et al. · May 21, 2026 Feucht et al. · May 14, 2026 Feucht et al. · May 14, 2026 Aranguri & Pernice · May 13, 2026 Aranguri & Pernice · May 13, 2026…
saved by
related reading
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Assessing skeptical views of interpretability research | Christopher Pottsweb.stanford.edu
- Understanding, Learning From, and Designing AI: Our Series Bgoodfire.ai
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- On Optimism for Interpretabilitygoodfire.ai
- Goodfire AIgoodfire.ai
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Research Areas in Interpretability (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org