Tsinghua paper: Does RL Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? — LessWrong
arXiv | project page | Authors: Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Yang Yue, Shiji Song, Gao Huang This paper from Tsinghua find that RL on verifiable rewards (RLVR) just increases the frequency at which capabilities are sampled, rather than giving a base model new capabilities. To do this, they compare pass@k scores between a base model and an RLed model. Recall that pass@k is the percentage of questions a model can solve at least once given k attempts at each question. Main result: On a math benchmark, an RLed model (yellow) has much better raw score / pass@1 than the base model (black), but lower pass@256! The authors say that RL prunes away reasoning pathways from the base model, but sometimes reasoning pathways that are rarely sampled end up being useful for solving the problem. So RL “narrows the reasoning boundary”— the region of problems the model is capable of solving sometimes. Thanks to @Vladimir_Nesov for mentioning this paper here. Excepting artificia
x Tsinghua paper: Does RL Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? — LessWrong Language Models (LLMs) AI Frontpage 70 Tsinghua paper: Does RL Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? by Thomas Kwa 5th May 2025 3 min read 22 70 This is a linkpost for https://arxiv.org/abs/2504.13837 arXiv | project page | Authors: Yang Yue , Zhiqi Chen , Rui Lu , Andrew Zhao , Zhaokai Wang , Yang Yue , Shiji Song , Gao Huang This paper from Tsinghua find that RL on verifiable rewards (RLVR) just increases the frequency at which capabilities are sampled, ra
Explore this link on the map →related reading
- Limit of RLVRlimit-of-rlvr.github.io
- DeepSeek-R1arxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Explore | alphaXivalphaxiv.org
- The State of Reinforcement Learning for LLM Reasoningsebastianraschka.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Understanding R1-Zero-Like Training: A Critical Perspectivearxiv.org
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com