CVAIJun 22

VideoLatent: Video-Language Learning via Latent Self-Forcing

arXiv:2606.2287024.0
Predicted impact top 6% in CV · last 90 daysOriginality Highly original
AI Analysis

This work addresses the high cost of CoT annotations and inference in video MLLMs, offering a scalable and efficient alternative for video understanding and reasoning tasks.

VideoLatent introduces a latent self-forcing training paradigm for video-language learning that eliminates the need for labor-intensive chain-of-thought annotations, achieving superior computational efficiency (6x training and 68x inference speedup over Video-R1) while outperforming existing models on 14 video understanding and reasoning benchmarks.

Recent advancements in chain-of-thought (CoT) reasoning have shown promise in enhancing video understanding and reasoning capabilities of multimodal large language models (MLLMs). However, existing CoT-based MLLMs require labor-intensive CoT annotations and incur substantial training and inference overhead. While visual latent reasoning has emerged as a more efficient alternative, existing methods primarily focus on image tasks and heavily rely on additional supervision signals for visual latent generation (e.g., CoT traces, auxiliary images, or fine-grained annotations), limiting their scalability and transferability to video tasks. To bridge this gap, we introduce VideoLatent, a novel MLLM equipped with a latent injection module tailored for video understanding and reasoning. Specifically, VideoLatent learns to perform visual latent reasoning using a new latent self-forcing training paradigm, which comprises latent alignment and latent diversity objectives, and relies solely on standard video-question-answer triplets. Extensive experiments across 14 benchmarks demonstrate that our model consistently outperforms existing standard and latent MLLMs on general video understanding and complex video reasoning. Compared with Video-R1, our VideoLatent achieves superior computational efficiency, reducing training/inference overhead by $\sim$6$\times$/$\sim$68$\times$. Moreover, experiments demonstrate that our method has strong generalizability to different MLLM backbones and different model scales.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes