VRPRM: Process Reward Modeling via Visual Reasoning
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
VRPRM: Process Reward Modeling via Visual Reasoning Xinquan Chen 1 , Bangwei Liu 1,2 , Xuhong Wang 1 , Yingchun Wang 1 , Chaochao Lu 1 Corresponding Author: wangxuhong@pjlab.org.cn Abstract Process Reward Model (PRM) is widely used in the post-training of Large Language Model (LLM) because it can perform fine-grained evaluation of the reasoning steps of generated content. However, most PRMs lack long-term reasoning and deep thinking capabilities. On the other hand, although a few works have tried to introduce Chain-of-Thought capability into PRMs, the annotation cost of CoT-PRM data is too exp
Explore this link on the map →saved by
related reading
- o1 and Reasoning | AndoLogsblog.ando.ai
- MolmoAct Action Reasoning Models that can Reason in Spacearxiv.org
- DeepSeek-R1arxiv.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Explore | alphaXivalphaxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- The State of Reinforcement Learning for LLM Reasoningsebastianraschka.com
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- RLHF Bookrlhfbook.com
- Limit of RLVRlimit-of-rlvr.github.io
- [2506.10947] Spurious Rewards: Rethinking Training Signals in RLVRarxiv.org