In (highly contingent!) defense of interpretability-in-the-loop ML training — AI Alignment Forum
Let’s call “interpretability-in-the-loop training” the idea of running a learning algorithm that involves an inscrutable trained model, and there’s some kind of interpretability system feeding into the loss function / reward function. Interpretability-in-the-loop training has a very bad rap (and rightly so). Here’s Yudkowsky 2022: When you explicitly optimize against a detector of unaligned thoughts, you're partially optimizing for more aligned thoughts, and partially optimizing for unaligned thoughts that are harder to detect. Optimizing against an interpreted thought optimizes against interpretability. Or Zvi 2025: The Most Forbidden Technique is training an AI using interpretability techniques. An AI produces a final output [X] via some method [M]. You can analyze [M] using technique [T], to learn what the AI is up to. You could train on that. Never do that. You train on [X]. Only [X]. Never [M], never [T]. Why? Because [T] is how you figure out when the model is misbehaving. If yo
x In (highly contingent!) defense of interpretability-in-the-loop ML training — AI Alignment Forum Agent Foundations Interpretability (ML & AI) AI Frontpage 35 In (highly contingent!) defense of interpretability-in-the-loop ML training by Steven Byrnes 6th Feb 2026 4 min read 11 35 Let’s call “interpretability-in-the-loop training” the idea of running a learning algorithm that involves an inscrutable trained model, and there’s some kind of interpretability system feeding into the loss function / reward function. Interpretability-in-the-loop training has a very bad rap (and rightly so). Here’s
saved by
related reading
- In (highly contingent!) defense of interpretability-in-the-loop ML training — LessWronglesswrong.com
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com
- Intentionally Designing the Future of AIgoodfire.ai
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org