flâneur — a map of the web's best reading

In (highly contingent!) defense of interpretability-in-the-loop ML training — AI Alignment Forum

alignmentforum.org · 2,120 words · saved by 2 readers

Let’s call “interpretability-in-the-loop training” the idea of running a learning algorithm that involves an inscrutable trained model, and there’s some kind of interpretability system feeding into the loss function / reward function. Interpretability-in-the-loop training has a very bad rap (and rightly so). Here’s Yudkowsky 2022: When you explicitly optimize against a detector of unaligned thoughts, you're partially optimizing for more aligned thoughts, and partially optimizing for unaligned thoughts that are harder to detect.  Optimizing against an interpreted thought optimizes against interpretability. Or Zvi 2025: The Most Forbidden Technique is training an AI using interpretability techniques. An AI produces a final output [X] via some method [M]. You can analyze [M] using technique [T], to learn what the AI is up to. You could train on that. Never do that. You train on [X]. Only [X]. Never [M], never [T]. Why? Because [T] is how you figure out when the model is misbehaving. If yo

x In (highly contingent!) defense of interpretability-in-the-loop ML training — AI Alignment Forum Agent Foundations Interpretability (ML & AI) AI Frontpage 35 In (highly contingent!) defense of interpretability-in-the-loop ML training by Steven Byrnes 6th Feb 2026 4 min read 11 35 Let’s call “interpretability-in-the-loop training” the idea of running a learning algorithm that involves an inscrutable trained model, and there’s some kind of interpretability system feeding into the loss function / reward function. Interpretability-in-the-loop training has a very bad rap (and rightly so). Here’s

Explore this link on the map →

saved by

related reading