CVAIJun 28

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos

arXiv:2606.2953119.0
Predicted impact top 8% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the need for fine-grained, region-specific motion understanding in videos, providing a benchmark and method for evaluating and improving video-language models on motion-centric tasks.

MotionAtlas introduces a system for region-aware motion captioning in videos, including a benchmark with 2,073 multiple-choice questions and a pipeline producing 159k training samples. Their model, MotionAtlas-4B, outperforms Qwen3-VL-4B by 5.2 percentage points on average across general motion benchmarks.

We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a family of powerful Video-MLLMs. Unlike conventional global motion captioning datasets, we focus on region-aware motion captioning: given a video and a spatiotemporal mask, the model generates precise descriptions of motion within the target region, thereby alleviating visual clutter and motion entanglement and enabling reliable, quantifiable evaluation. Concretely, we first build MotionAtlas-Bench, a comprehensive benchmark comprising 2,073 multiple-choice questions, meticulously annotated for a curated set of high-quality, motion-centric videos, to evaluate fine-grained motion understanding of the objects in question. Second, we design a rigorous and scalable data pipeline that leverages self-bootstrap refinement to suppress fine-grained hallucinations, yielding 159k high-quality motion captioning data. Third, we design a tailored training data composition strategy, which achieves consistent and substantial performance gains across diverse baseline Video-MLLMs, including Molmo2 and Qwen3-VL. For instance, MotionAtlas-4B surpasses Qwen3-VL-4B by an average of 5.2 percentage points across general motion benchmarks. The benchmark, dataset, and code have been released.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes