CVAIJun 29

SUMO: Segment and Track Any Motion with Nonlinear State Space Models

arXiv:2606.298617.8
Predicted impact top 55% in CV · last 90 daysOriginality Highly original
AI Analysis

For computer vision tasks requiring robust tracking and segmentation in complex motion scenarios, SUMO addresses the limitation of existing methods that rely solely on visual cues.

SUMO proposes a zero-shot, training-free framework that integrates nonlinear dynamics with vision-based segmentation to improve Visual Object Tracking and Moving Object Segmentation. It achieves state-of-the-art performance on both tasks.

Visual Object Tracking (VOT) and Moving Object Segmentation (MOS) are two fundamental tasks in computer vision that involve both spatial and temporal object dynamics. Existing methods rely predominantly on visual cues and thus often falter in real-world scenarios where object motions are inherently complex and nonlinear. To address this limitation, we propose SUMO, a zero-shot, training-free, unified framework integrating nonlinear dynamics with vision-based segmentation for accurate and consistent VOT and MOS. Specifically, we develop a nonlinear State Space Model (SSM) inspired by robotics principles to capture the complex object dynamics. Building on this model, we propose a Selective Unscented Filter (SUF) for accurate state estimation, which features a joint scoring mechanism and dynamically fuses multi-source predictions to identify the most plausible object state over time. Furthermore, we apply a memory selection mechanism to evaluate the reliability of memory frames. Our extensive experimental results show that SUMO achieves state-of-the-art performance on both VOT and MOS tasks.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes