flâneur — a map of the web's best reading

RL with Spurious Rewards

rethink-rlvr.notion.site · 279 words · saved by 5 readers

💭 Spurious Rewards: Rethinking Training Signals in RLVR Get Notion free 💭 Spurious Rewards: Rethinking Training Signals in RLVR Rulin Shao*, Shuyue Stella Li*, Rui Xin*, Scott Geng*, Yiping Wang, Sewoong Oh, Simon Shaolei Du, Nathan Lambert, Sewon Min, Ranjay Krishna, Yulia Tsvetkov, Hannaneh Hajishirzi, Pang Wei Koh, Luke Zettlemoyer *Equal Contribution Paper: paper/rethink-rlvr.pdfruixin31/Rethink_RLVR​ Github: Rethink_RLVR​ 💡 TL;DR We show that you can do RLVR on Qwen2.5-Math models with completely random or incorrect rewards, and still get massive math benchmark gains. All of the following spurious rewards give 15-20+ points on MATH-500 when RLVR training Qwen2.5-Math-7B: RLVR + format reward (reward responses with \boxed{}) 🔲: +16.4% RLVR + incorrect reward (only incorrect answers rewarded) 😈: +24.6% RLVR + random reward 🎲: +21.4% (as a reference) RLVR + ground-truth reward ✅: + 28.8% 🤯 How can these spurious rewards possibly work? Can we get similar gains on other mode

Explore this link on the map →

saved by