CVJul 3

Probing Identity-Specific Motion Signatures: A Controlled Diagnostic Study

arXiv:2607.0363312.5
Predicted impact top 28% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in video understanding and re-identification, this work reveals that current models overlook motion signatures in favor of static shortcuts, highlighting the need for methods that leverage motion.

This study investigates whether modern video models use identity-specific motion cues for recognition, finding that they rely on static appearance cues (faces, jerseys) even when motion is informative. When appearance is suppressed (silhouette/skeleton inputs), models shift to motion micro-patterns, achieving competitive accuracy and stronger robustness to appearance shifts.

Identity recognition (e.g., person, animal re-identification) has traditionally relied heavily on static appearance cues. Yet motion--consistent, individual-specific dynamics--can provide a complementary and potentially more robust signature, especially when appearance is weak or variable. This raises a fundamental question: when identity-specific motion cues are clearly present, to what extent do modern video models use them for recognition? To investigate this question, we conduct a systematic diagnostic study and introduce BALLER120, a controlled benchmark of 120 professional basketball players performing free-throws. By focusing on the same multi-phase action across individuals, BALLER120 reduces action-level variation and identity-correlated acquisition biases, enabling fine-grained analysis of identity-specific kinematic patterns. We find that modern video models can predict identity accurately from RGB videos, but often rely on static appearance cues such as faces and jersey regions, even when informative motion cues are available. Strikingly, when appearance is suppressed through silhouette-only or skeleton-only inputs, the same model architectures shift toward motion micro-patterns (e.g., foot placement and elbow bending). Despite containing less visual information, appearance-suppressed representations achieve competitive accuracy and stronger robustness to appearance shifts. Our qualitative analyses further show that appearance-suppressed models attend to distinctive motion patterns across individuals. Overall, our study demonstrates that identity-specific motion signatures are present, informative, and learnable, but modern video models may overlook them in favor of easier static shortcuts unless appearance cues are explicitly suppressed.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes