Limit of RLVR
Yang Yue is currently focused on developing new paradigms for incentivizing LLM/MLLM reasoning, generalized world models, and exploring the generalization of VLA. He is seeking active collaboration opportunities with companies that offer the freedom to explore these frontier and fundamental questions, alongside abundant resources and a strong technical atmosphere. Additionally, he is seeking a Ph.D. visit. Please feel free to reach out if there is potential for collaboration. Video: pass@k curves of base models and their zero-RL-trained counterparts across multiple mathematical benchmarks. When k is small, RL-trained models outperform their base versions. However, as k increases to the tens or hundreds, base models consistently catch up with RL-trained models across all benchmarks and LLM families without exception. Eventually, base models surpass RL-trained models. Recent breakthroughs in reasoning-focused large language models (LLMs) like OpenAI-o1, DeepSeek-R1, and Kimi-1.5 have lar
Limit of RLVR Limit of RLVR Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? We systematically study Reinforcement Learning with Verifiable Rewards (RLVR) across math, coding, and vision benchmarks and uncover that RL fine-tuning enhances sampling efficiency without expanding the reasoning capacity already present in base models. Base models surpass RL at large pass@k. RL-trained variants shine at low sampling budgets, yet the original base models consistently overtake them as k increases. RL narrows exploration. rewarded trajectories are amplifi
related reading
- [2504.13837] Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?arxiv.org
- Tsinghua paper: Does RL Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? — LessWronglesswrong.com
- DeepSeek-R1arxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- As Rocks May Think | Eric Jangevjang.com
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- The Invisible Leash: Why RLVR May Not Escape Its Originarxiv.org
- Reasoning in General Domains without Verifiersarxiv.org
- LLM Reasoning via One Examplearxiv.org
- Explore | alphaXivalphaxiv.org
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectoriesarxiv.org