flâneur

Models know when they’re reward hacking — and we can catch them at scale - Goodfire

goodfire.com · saved by 1 readers

We found a clear internal signal in models that accompanies reward hacking, and built probes that detect it — enabling efficient, real-time detection of reward hacking at scale.

saved by