flâneur — a map of the web's best reading

Limit of RLVR

limit-of-rlvr.github.io · 2,804 words · saved by 1 readers

Yang Yue is currently focused on developing new paradigms for incentivizing LLM/MLLM reasoning, generalized world models, and exploring the generalization of VLA. He is seeking active collaboration opportunities with companies that offer the freedom to explore these frontier and fundamental questions, alongside abundant resources and a strong technical atmosphere. Additionally, he is seeking a Ph.D. visit. Please feel free to reach out if there is potential for collaboration. Video: pass@k curves of base models and their zero-RL-trained counterparts across multiple mathematical benchmarks. When k is small, RL-trained models outperform their base versions. However, as k increases to the tens or hundreds, base models consistently catch up with RL-trained models across all benchmarks and LLM families without exception. Eventually, base models surpass RL-trained models. Recent breakthroughs in reasoning-focused large language models (LLMs) like OpenAI-o1, DeepSeek-R1, and Kimi-1.5 have lar

Limit of RLVR Limit of RLVR Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? We systematically study Reinforcement Learning with Verifiable Rewards (RLVR) across math, coding, and vision benchmarks and uncover that RL fine-tuning enhances sampling efficiency without expanding the reasoning capacity already present in base models. Base models surpass RL at large pass@k. RL-trained variants shine at low sampling budgets, yet the original base models consistently overtake them as k increases. RL narrows exploration. rewarded trajectories are amplifi

Explore this link on the map →

related reading