CVSDJul 1

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization

arXiv:2607.007267.7
Predicted impact top 56% in CV · last 90 daysOriginality Incremental advance
AI Analysis

Provides a decoupled evaluation framework for audio-visual synchronization, addressing a key limitation in current benchmarks for multimodal understanding and generation tasks.

Existing audio-visual synchronization benchmarks conflate temporal and semantic evaluation. AV-SyncBench is the first benchmark to fully decouple these dimensions, spanning 3,269 videos and 38,390 samples across 10 scenarios, enabling independent assessment of temporal and semantic consistency.

Audio-visual feature extraction is a fundamental component of multimodal understanding and generation tasks. However, existing evaluation protocols for feature extraction models exhibit dimensional bias, typically focusing on either semantic matching or temporal offset detection. Moreover, their data construction remains coupled, preventing independent assessment of temporal and semantic consistency. We propose AV-SyncBench, the first benchmark to fully separate temporal and semantic evaluation for audio-visual synchronization. Built from in-the-wild videos, it spans Voice, Music, and Sound across 10 scenarios and 5 challenge tasks. Data are automatically filtered and manually verified to ensure on-screen sound sources. The benchmark contains 3,269 videos and 38,390 samples, and we evaluate five representative models to quantify feature quality for alignment and downstream tasks. The code and dataset are available at: https://fgt7t6g.github.io/AV-SyncBench.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes