flâneur — a map of the web's best reading

Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWrong

lesswrong.com · 5,068 words · saved by 1 readers

Something’s changed about reward hacking[1] in recent systems. In the past, reward hacks were usually accidents, found by non-general, RL-trained sys…

x Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWrong Language Models (LLMs) Reinforcement learning Reward Functions AI Frontpage 97 Reward hacking is becoming more sophisticated and deliberate in frontier LLMs by Kei Nishimura-Gasparian 24th Apr 2025 1 min read 7 97 Something’s changed about reward hacking [1] in recent systems. In the past, reward hacks were usually accidents, found by non-general, RL-trained systems. Models would randomly explore different behaviors and would sometimes come across undesired behaviors that achieved high rewards [2] . The

Explore this link on the map →

saved by

related reading