flâneur — a map of the web's best reading

Features as Rewards: Using Interpretability to Reduce Hallucinations

goodfire.ai · 2,191 words · saved by 2 readers

In our recent essay on intentional design, we described a vision for using interpretability to guide model training — moving from guess-and-check to closed-loop control. Our new paper introduces our first concrete demonstration of that vision: using lightweight probes on a model's internal representations as reward signals for reinforcement learning, applied to the problem of reducing hallucinations. We call this approach RLFR: Reinforcement Learning from Feature Rewards. Our results show that RLFR is both effective in terms of reducing hallucinations without off-target effects, and retains the ability to use the probe as a monitor at test time, which unlocks strong test-time scaling performance. Quantitatively, RLFR reduces hallucinations in Gemma-3-12B-IT by 58% (when run with our probing harness), at ~90× lower cost per intervention than the LLM-as-judge alternative, with no degradation on standard benchmarks. You can explore some samples in our interactive viewer here. Teaching a l

Features as Rewards: Using Interpretability to Reduce Hallucinations Research Features as Rewards: Using Interpretability to Reduce Hallucinations Authors Aaditya Prasad * † Connor Watts * † Jack Merullo † Dhruvil Gala † Owen Lewis † Thomas McGrath † Ekdeep Singh Lubana † * Equal contribution † Goodfire Blog post by Tom McGrath and Michael Byun Published February 11, 2026 Full Paper Read on arXiv → Your browser does not support the video tag. Contents Why is it hard to fix hallucinations? Our approach Probe pipeline: detecting hallucinations in complex long-form responses RL setup Results: red

Explore this link on the map →

saved by

related reading