CVAIJul 16

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

arXiv:2607.1471110.3h-index: 4
Predicted impact top 38% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For video understanding researchers, VideoSEMA offers a scalable and efficient alternative to existing attention mechanisms, outperforming larger models on standard benchmarks.

VideoSEMA introduces a split space-time attention model combining a Mamba-like attention block for spatial processing with softmax temporal attention, achieving superior accuracy on K400 and SSv2 benchmarks compared to heavier ViT and Mamba models, while maintaining graceful accuracy degradation at higher resolutions (e.g., 1024^2) without fine-tuning.

We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time. In each frame, SEMA attention applies a local window attention in parallel with a global averaging in a Mamba macro-architecture, which is called Mamba-like. Under certain rank conditions, we prove that the computationally cheaper split space-time attention is equivalent to full space-time attention. On benchmark K400 data sets, VideoSEMA out-performs heavier vision transformer and Mamba models. On benchmark SSv2 data, VideoSEMA leads in top-1 accuracy among models of similar parameter sizes. As image resolution scales up from standard $224^2$ to $1024^2$ on K400 and without fine-tuning, VideoSEMA degrades much more gracefully than VideoMamba in accuracy. It is promising to extend VideoSEMA to longer videos with a dilated/sparse temporal attention.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes