In (highly contingent!) defense of interpretability-in-the-loop ML training — LessWrong
Let’s call “interpretability-in-the-loop training” the idea of running a learning algorithm that involves an inscrutable trained model, and there’s s…
x In (highly contingent!) defense of interpretability-in-the-loop ML training — LessWrong Agent Foundations Interpretability (ML & AI) AI Frontpage 85 In (highly contingent!) defense of interpretability-in-the-loop ML training by Steven Byrnes 6th Feb 2026 AI Alignment Forum 4 min read 11 85 Ω 35 Let’s call “interpretability-in-the-loop training” the idea of running a learning algorithm that involves an inscrutable trained model, and there’s some kind of interpretability system feeding into the loss function / reward function. Interpretability-in-the-loop training has a very bad rap (and right
related reading
- In (highly contingent!) defense of interpretability-in-the-loop ML training — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- Intentionally Designing the Future of AIgoodfire.ai
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org