Features as Rewards: Using Interpretability to Reduce Hallucinations
In our recent essay on intentional design, we described a vision for using interpretability to guide model training — moving from guess-and-check to closed-loop control. Our new paper introduces our first concrete demonstration of that vision: using lightweight probes on a model's internal representations as reward signals for reinforcement learning, applied to the problem of reducing hallucinations. We call this approach RLFR: Reinforcement Learning from Feature Rewards. Our results show that RLFR is both effective in terms of reducing hallucinations without off-target effects, and retains the ability to use the probe as a monitor at test time, which unlocks strong test-time scaling performance. Quantitatively, RLFR reduces hallucinations in Gemma-3-12B-IT by 58% (when run with our probing harness), at ~90× lower cost per intervention than the LLM-as-judge alternative, with no degradation on standard benchmarks. You can explore some samples in our interactive viewer here. Teaching a l
Features as Rewards: Using Interpretability to Reduce Hallucinations Research Features as Rewards: Using Interpretability to Reduce Hallucinations Authors Aaditya Prasad * † Connor Watts * † Jack Merullo † Dhruvil Gala † Owen Lewis † Thomas McGrath † Ekdeep Singh Lubana † * Equal contribution † Goodfire Blog post by Tom McGrath and Michael Byun Published February 11, 2026 Full Paper Read on arXiv → Your browser does not support the video tag. Contents Why is it hard to fix hallucinations? Our approach Probe pipeline: detecting hallucinations in complex long-form responses RL setup Results: red
Explore this link on the map →saved by
related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- DeepSeek-R1arxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- In (highly contingent!) defense of interpretability-in-the-loop ML training — AI Alignment Forumalignmentforum.org
- How confessions can keep language models honest | OpenAIopenai.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com