CVHCJun 30

AA: A Multi-view Multimodal Dataset for Screen-based Gaze Estimation

arXiv:2606.312118.81 citations
Predicted impact top 47% in CV · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers in gaze estimation, this dataset addresses the lack of multi-view data, but it is an incremental contribution as it primarily provides a new data resource without demonstrating novel methods or significant performance gains.

AA introduces a multi-view multimodal dataset for screen-based gaze estimation, featuring synchronized facial observations from multiple cameras and precise gaze targets, enabling more robust modeling under viewpoint variation and occlusion.

We present AA, a multi-view multimodal dataset for screen-based gaze estimation. The dataset captures synchronized facial observations from eight fixed screen-mounted cameras and two additional side-view cameras, paired with precise screen-space gaze targets collected under controlled fixation conditions. Each sample contains multi-view face observations together with structured facial region crops, enabling multimodal learning from both global and local visual cues. Unlike existing single-view gaze datasets, AA provides multi-view coverage from both screen-mounted and side-mounted perspectives, enabling more robust modeling under viewpoint variation and occlusion. The dataset includes subject-independent evaluation splits and a standardized data processing pipeline to support reproducible research in gaze estimation.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes