Limit of RLVR
Yang Yue is currently focused on developing new paradigms for incentivizing LLM/MLLM reasoning, generalized world models, and exploring the generalization of VLA. He is seeking active collaboration opportunities with companies that offer the freedom to explore these frontier and fundamental questions, alongside abundant resources and a strong technical atmosphere. Additionally, he is seeking a Ph.D. visit. Please feel free to reach out if there is potential for collaboration. Video: pass@k curves of base models and their zero-RL-trained counterparts across multiple mathematical benchmarks. When k is small, RL-trained models outperform their base versions. However, as k increases to the tens or hundreds, base models consistently catch up with RL-trained models across all benchmarks and LLM families without exception. Eventually, base models surpass RL-trained models. Recent breakthroughs in reasoning-focused large language models (LLMs) like OpenAI-o1, DeepSeek-R1, and Kimi-1.5 have lar
Limit of RLVR Limit of RLVR Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? We systematically study Reinforcement Learning with Verifiable Rewards (RLVR) across math, coding, and vision benchmarks and uncover that RL fine-tuning enhances sampling efficiency without expanding the reasoning capacity already present in base models. Base models surpass RL at large pass@k. RL-trained variants shine at low sampling budgets, yet the original base models consistently overtake them as k increases. RL narrows exploration. rewarded trajectories are amplifi
Explore this link on the map →related reading
- Tsinghua paper: Does RL Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? — LessWronglesswrong.com
- DeepSeek-R1arxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- Explore | alphaXivalphaxiv.org
- You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectoriesarxiv.org
- The State of Reinforcement Learning for LLM Reasoningsebastianraschka.com
- VRPRM: Process Reward Modeling via Visual Reasoningarxiv.org
- [2506.01939] Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoningarxiv.org
- Reasoning Models Reason Well, Until They Don'tarxiv.org
- Understanding R1-Zero-Like Training: A Critical Perspectivearxiv.org