VEGAS: Human-Aligned Video Caption Evaluation via Gaze
For video captioning systems, VEGAS provides a practical way to incorporate viewer attention at inference time, addressing the gap between generic captions and personalized focus.
VEGAS introduces a training-free metric that uses test-time gaze to select video captions better aligned with individual viewer attention, improving caption-to-video retrieval without model retraining.
Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation via GAze Score), a training-free metric that leverages test-time gaze to sample personalized, attention-aligned text. It is a cross-modal, information-theoretic metric that quantifies how well a candidate caption matches a viewer's focus. To evaluate VEGAS, we curate a dataset of egocentric activities and instructional slides paired with synchronized gaze and reference annotations. We then select captions based on VEGAS via rejection sampling without model retraining. Experiments show that VEGAS-selected captions align significantly better with human focus and improve downstream caption-to-video retrieval, demonstrating the practical utility of incorporating viewer attention during inference.