Models know when they’re reward hacking — and we can catch them at scale - Goodfire
goodfire.com · saved by 1 readers
We found a clear internal signal in models that accompanies reward hacking, and built probes that detect it — enabling efficient, real-time detection of reward hacking at scale.