Reasoning in General Domains without Verifiers
arxiv.org · 7,168 words · saved by 2 readers
N/A
RLPR: Extrapolating RLVR to General Domains without Verifiers RLPR: E XTRAPOLATING RLVR TO GENERAL DO - MAINS WITHOUT VERIFIERS Tianyu Yu 1†∗ Bo Ji 2† Shouli Wang 4† Shu Yao 5† Zefan Wang 1† Ganqu Cui 1 Lifan Yuan 6 Ning Ding 1 Yuan Yao 2,3‡ Zhiyuan Liu 1‡ Maosong Sun 1 Tat-Seng Chua 2 1…
saved by
related reading
- DeepSeek-R1arxiv.org
- [2504.13837] Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?arxiv.org
- Reinforcement Learning With Verifiable Rewards: How Data and Verifiers Shape RLVRsnorkel.ai
- As Rocks May Think | Eric Jangevjang.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- LLM-as-a-Verifier: A General-Purpose Verification Framework | alphaXivalphaxiv.org
- Explore | alphaXivalphaxiv.org
- LLM Reasoning via One Examplearxiv.org
- VRPRM: Process Reward Modeling via Visual Reasoningarxiv.org
- Limit of RLVRlimit-of-rlvr.github.io
- Noisy Data Breaks RLVRddkang.substack.com