CVJul 2

LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension

arXiv:2607.0209612.9
Predicted impact top 25% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This benchmark addresses the lack of long-form, untrimmed egocentric video datasets for referring expression comprehension, exposing limitations of existing models for real-world applications.

The paper introduces LongEgoRefer, a benchmark for long-form egocentric video referring expression comprehension, featuring 1,498 queries over 45-minute videos with sparse target occurrences. Current state-of-the-art models struggle significantly on this benchmark, highlighting the difficulty of long-form spatio-temporal grounding.

Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this context, Video Referring Expression Comprehension (Video REC), the task of localizing the temporal and spatial extent of a referred object in video frames given a natural language query, plays a key role in linking textual descriptions to observed objects in untrimmed egocentric recordings. However, existing egocentric Video REC benchmarks primarily focus on short video clips, where some target object appears densely within frames. Such settings do not reflect real-world egocentric recordings, which are long-form, untrimmed, and characterized by sparse object occurrences and complex activity transitions. To address this limitation, we introduce LongEgoRefer, a novel and challenging benchmark constructed from long-form videos in the Ego4D dataset. LongEgoRefer contains 1,498 referring expressions with an average video duration of 45 minutes. The benchmark exhibits extreme target sparsity, detailed linguistic descriptions, and complex human-object interactions embedded in long, dynamic egocentric narratives. Consequently, it defines a demanding spatio-temporal grounding problem that requires models to identify both when an event occurs and where the referred object appears within extended video sequences. We evaluate existing Video REC approaches, including training-free baselines based on vision-language models combined with Grounded SAM2. Extensive experiments show that even advanced baselines and current state-of-the-art models struggle significantly on LongEgoRefer. These results highlight the intrinsic difficulty of long-form egocentric spatio-temporal grounding and emphasize the need for more robust video understanding models.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes