Noisy Data Breaks RLVR - by Daniel Kang - Daniel’s Substack
ddkang.substack.com · 782 words · saved by 1 readers
Reinforcement learning with verifiable rewards (RLVR) is a widely used post-training paradigm to improve the reasoning capabilities of LLMs.
Reinforcement learning with verifiable rewards (RLVR) is a widely used post-training paradigm to improve the reasoning capabilities of LLMs. However, creating the high-quality verifiable answers needed to train the model is labor-intensive and expensive, as highlighted by the boom of data labeling companies, such as Mercor. This creates a fundamental question for post-training: To what extent can RLVR, with algorithmic improvements, tolerate large-volume, noisy data? Recent literature promotes a counter-intuitive idea: RLVR is robust to noisy data. Prior work shows that training on “100%…
saved by
related reading
- Reinforcement Learning With Verifiable Rewards: How Data and Verifiers Shape RLVRsnorkel.ai
- State of RL for reasoning LLMs | A. Weersaweers.de
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- Reasoning in General Domains without Verifiersarxiv.org
- DeepSeek-R1arxiv.org
- [2504.13837] Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?arxiv.org
- Towards a Typology of Strange LLM Chains-of-Thought1a3orn.com
- LLM Reasoning via One Examplearxiv.org
- The Invisible Leash: Why RLVR May Not Escape Its Originarxiv.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- [2506.10947] Spurious Rewards: Rethinking Training Signals in RLVRarxiv.org