flâneur

Noisy Data Breaks RLVR - by Daniel Kang - Daniel’s Substack

ddkang.substack.com · 782 words · saved by 1 readers

Reinforcement learning with verifiable rewards (RLVR) is a widely used post-training paradigm to improve the reasoning capabilities of LLMs.

Reinforcement learning with verifiable rewards (RLVR) is a widely used post-training paradigm to improve the reasoning capabilities of LLMs. However, creating the high-quality verifiable answers needed to train the model is labor-intensive and expensive, as highlighted by the boom of data labeling companies, such as Mercor. This creates a fundamental question for post-training: To what extent can RLVR, with algorithmic improvements, tolerate large-volume, noisy data? Recent literature promotes a counter-intuitive idea: RLVR is robust to noisy data. Prior work shows that training on “100%…

saved by

related reading