flâneur

Claim-Level Rubric Rewards for Video Caption Reinforcement Learning

arxiv.org · 8,384 words · saved by 1 readers

In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense video captioning. Existing reward designs generally fall into two categories: holistic response-level judgment across heterogeneous criteria, or alignment-based evaluation against reference captions. However, both paradigms suffer from fundamental limitations. Holistic rewards struggle to ensure factual accuracy and are prone to stylistic reward hacking, while reference-based rewards overly rely on rigid textual alignment, failing to preserve the completeness and diversity inherent to open-ended generation tasks. To address these challenges, CuRe reformulates reward modeling as fine-grained claim-level verification. Specifically, CuRe decomposes captions into category-aware atomic claims through a structured rubric, converting holistic evaluation into simpler and more reliable claim-level verification. To rewar

Mingqi Gao1,3,∗, Hongyuan Dong3,∗,†, Yifei Chen3,∗, Zhisheng Zhong3, Zheng Ruan3, Wenjin Hou3, Yu Chen2, Han Hu3,‡, Yansong Tang1,🖂{}^{1,\text{\Letter}} Affiliation: 1Tsinghua Shenzhen International Graduate School, Tsinghua University 2University of Chinese Academy of Sciences Affiliation: 3LLM Department, Tencent Email: minkkigao@gmail.com tang.yansong@sz.tsinghua.edu.cn Abstract In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense video captioning. Existing…

saved by

related reading