Renormalization Redux: QFT Techniques for AI Interpretability — AI Alignment Forum
In a previous post, Lauren offered a take on why a physics way of thinking is so successful at understanding AI systems. In this post, we look in more detail at the potential of Quantum field theory (QFT) to be expanded into a more comprehensive framework for this purpose. Interest in this area has been steadily increasing[1], but efforts have yet to condense into a larger-scale, coordinated effort. In particular, a lot of the more theoretical, technically detailed work remains opaque to anyone not well-versed in physics, meaning that insights[2] are largely disconnected from the AI safety community. The most accessible of these is Principles of Deep Learning theory (which we abbreviate “PDLT”), a nearly 500 page book that lays the groundwork for these ideas[3]. While there has been some AI safety research that has incorporated QFT-inspired threads[4], we see untapped potential for cross-disciplinary collaborations to unify these disparate directions. With this post – one of several in
x Renormalization Redux: QFT Techniques for AI Interpretability — AI Alignment Forum Renormalizing interpretability AI Frontpage 19 Renormalization Redux: QFT Techniques for AI Interpretability by Lauren Greenspan , Dmitry Vaintrob 18th Jan 2025 9 min read 12 19 Introduction: Why QFT? In a previous post , Lauren offered a take on why a physics way of thinking is so successful at understanding AI systems. In this post, we look in more detail at the potential of Quantum field theory (QFT) to be expanded into a more comprehensive framework for this purpose. Interest in this area has been steadily
Explore this link on the map →related reading
- Renormalization Redux: QFT Techniques for AI Interpretability — LessWronglesswrong.com
- Renormalization Roadmap — LessWronglesswrong.com
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- On neural scaling and the quanta hypothesisericjmichaud.com
- Statistical Physics for Ambitious Interpretability: A Workshop Retrospective — LessWronglesswrong.com
- On Optimism for Interpretabilitygoodfire.ai
- The Building Blocks of Interpretabilitydistill.pub
- On Developing a Mathematical Theory of Interpretability — LessWronglesswrong.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Problem Areas in Physics and AI Safety | Apart Researchapartresearch.com
- Intentionally Designing the Future of AIgoodfire.ai