Claim-Level Rubric Rewards for Video Caption Reinforcement Learning
In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense video captioning. Existing reward designs generally fall into two categories: holistic response-level judgment across heterogeneous criteria, or alignment-based evaluation against reference captions. However, both paradigms suffer from fundamental limitations. Holistic rewards struggle to ensure factual accuracy and are prone to stylistic reward hacking, while reference-based rewards overly rely on rigid textual alignment, failing to preserve the completeness and diversity inherent to open-ended generation tasks. To address these challenges, CuRe reformulates reward modeling as fine-grained claim-level verification. Specifically, CuRe decomposes captions into category-aware atomic claims through a structured rubric, converting holistic evaluation into simpler and more reliable claim-level verification. To rewar
Mingqi Gao1,3,∗, Hongyuan Dong3,∗,†, Yifei Chen3,∗, Zhisheng Zhong3, Zheng Ruan3, Wenjin Hou3, Yu Chen2, Han Hu3,‡, Yansong Tang1,🖂{}^{1,\text{\Letter}} Affiliation: 1Tsinghua Shenzhen International Graduate School, Tsinghua University 2University of Chinese Academy of Sciences Affiliation: 3LLM Department, Tencent Email: minkkigao@gmail.com tang.yansong@sz.tsinghua.edu.cn Abstract In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense video captioning. Existing…
saved by
related reading
- GitHub - FusionBrainLab/Vision_GRPOgithub.com
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domainsarxiv.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- DeepSeek-R1arxiv.org
- Part 1: The Map Was Wrong — Nemostationnemostation.com
- [2605.12474] Reward Hacking in Rubric-Based Reinforcement Learningarxiv.org
- 2310.12921.pdfarxiv.org
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedbackarxiv.org
- VRPRM: Process Reward Modeling via Visual Reasoningarxiv.org
- Measuring Reward-Seeking by Instilling Contrastive Beliefsalignment.openai.com