flâneur — a map of the web's best reading

Tsinghua paper: Does RL Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? — LessWrong

lesswrong.com · 4,253 words · saved by 1 readers

arXiv | project page | Authors: Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Yang Yue, Shiji Song, Gao Huang This paper from Tsinghua find that RL on verifiable rewards (RLVR) just increases the frequency at which capabilities are sampled, rather than giving a base model new capabilities. To do this, they compare pass@k scores between a base model and an RLed model. Recall that pass@k is the percentage of questions a model can solve at least once given k attempts at each question. Main result: On a math benchmark, an RLed model (yellow) has much better raw score / pass@1 than the base model (black), but lower pass@256! The authors say that RL prunes away reasoning pathways from the base model, but sometimes reasoning pathways that are rarely sampled end up being useful for solving the problem. So RL “narrows the reasoning boundary”— the region of problems the model is capable of solving sometimes. Thanks to @Vladimir_Nesov for mentioning this paper here. Excepting artificia

x Tsinghua paper: Does RL Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? — LessWrong Language Models (LLMs) AI Frontpage 70 Tsinghua paper: Does RL Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? by Thomas Kwa 5th May 2025 3 min read 22 70 This is a linkpost for https://arxiv.org/abs/2504.13837 arXiv | project page | Authors: Yang Yue , Zhiqi Chen , Rui Lu , Andrew Zhao , Zhaokai Wang , Yang Yue , Shiji Song , Gao Huang This paper from Tsinghua find that RL on verifiable rewards (RLVR) just increases the frequency at which capabilities are sampled, ra

Explore this link on the map →

related reading