In (highly contingent!) defense of interpretability-in-the-loop ML training — AI Alignment Forum
Let’s call “interpretability-in-the-loop training” the idea of running a learning algorithm that involves an inscrutable trained model, and there’s some kind of interpretability system feeding into the loss function / reward function. Interpretability-in-the-loop training has a very bad rap (and rightly so). Here’s Yudkowsky 2022: When you explicitly optimize against a detector of unaligned thoughts, you're partially optimizing for more aligned thoughts, and partially optimizing for unaligned thoughts that are harder to detect. Optimizing against an interpreted thought optimizes against interpretability. Or Zvi 2025: The Most Forbidden Technique is training an AI using interpretability techniques. An AI produces a final output [X] via some method [M]. You can analyze [M] using technique [T], to learn what the AI is up to. You could train on that. Never do that. You train on [X]. Only [X]. Never [M], never [T]. Why? Because [T] is how you figure out when the model is misbehaving. If yo
x In (highly contingent!) defense of interpretability-in-the-loop ML training — AI Alignment Forum Agent Foundations Interpretability (ML & AI) AI Frontpage 35 In (highly contingent!) defense of interpretability-in-the-loop ML training by Steven Byrnes 6th Feb 2026 4 min read 11 35 Let’s call “interpretability-in-the-loop training” the idea of running a learning algorithm that involves an inscrutable trained model, and there’s some kind of interpretability system feeding into the loss function / reward function. Interpretability-in-the-loop training has a very bad rap (and rightly so). Here’s
Explore this link on the map →saved by
related reading
- In (highly contingent!) defense of interpretability-in-the-loop ML training — LessWronglesswrong.com
- Intentionally Designing the Future of AIgoodfire.ai
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- On Optimism for Interpretabilitygoodfire.ai
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io