CVJul 6

Claim-Level Rubric Rewards for Video Caption Reinforcement Learning

arXiv:2607.0515019.4
Predicted impact top 9% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the reward-design bottleneck in reinforcement learning for dense video captioning, offering a more reliable and fine-grained evaluation method.

CuRe introduces a structured reward framework for dense video captioning that decomposes captions into category-aware atomic claims for fine-grained verification, achieving improved factual accuracy and diversity over holistic or reference-based rewards.

In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense video captioning. Existing reward designs generally fall into two categories: holistic response-level judgment across heterogeneous criteria, or alignment-based evaluation against reference captions. However, both paradigms suffer from fundamental limitations. Holistic rewards struggle to ensure factual accuracy and are prone to stylistic reward hacking, while reference-based rewards overly rely on rigid textual alignment, failing to preserve the completeness and diversity inherent to open-ended generation tasks. To address these challenges, CuRe reformulates reward modeling as fine-grained claim-level verification. Specifically, CuRe decomposes captions into category-aware atomic claims through a structured rubric, converting holistic evaluation into simpler and more reliable claim-level verification.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes