RL with Spurious Rewards
💭 Spurious Rewards: Rethinking Training Signals in RLVR Get Notion free 💭 Spurious Rewards: Rethinking Training Signals in RLVR Rulin Shao*, Shuyue Stella Li*, Rui Xin*, Scott Geng*, Yiping Wang, Sewoong Oh, Simon Shaolei Du, Nathan Lambert, Sewon Min, Ranjay Krishna, Yulia Tsvetkov, Hannaneh Hajishirzi, Pang Wei Koh, Luke Zettlemoyer *Equal Contribution Paper: paper/rethink-rlvr.pdfruixin31/Rethink_RLVR Github: Rethink_RLVR 💡 TL;DR We show that you can do RLVR on Qwen2.5-Math models with completely random or incorrect rewards, and still get massive math benchmark gains. All of the following spurious rewards give 15-20+ points on MATH-500 when RLVR training Qwen2.5-Math-7B: RLVR + format reward (reward responses with \boxed{}) 🔲: +16.4% RLVR + incorrect reward (only incorrect answers rewarded) 😈: +24.6% RLVR + random reward 🎲: +21.4% (as a reference) RLVR + ground-truth reward ✅: + 28.8% 🤯 How can these spurious rewards possibly work? Can we get similar gains on other mode
Explore this link on the map →