SDAIHCJun 15

MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild

arXiv:2606.167314.5
Predicted impact top 81% in SD · last 90 daysOriginality Incremental advance
AI Analysis

Enables practical turn-taking prediction for human-robot interaction without complex sensor setups, addressing a key bottleneck in deploying conversational AI in the wild.

MuVAP introduces a causal multimodal framework for turn-taking prediction in multiparty conversations using only monaural audio and a single camera, outperforming baselines on Shift-Hold and next-speaker prediction tasks across two- and three-speaker settings.

Current multiparty turn-taking models often rely on complex microphone arrays or multi-camera setups, limiting their applicability in human-robot interaction scenarios. We introduce MuVAP, a causal multimodal framework that extends Voice Activity Projection by grounding acoustic predictions in face tracks, enabling speaker-aware turn-taking predictions from a monaural audio stream and a single camera view. To address the combinatorial complexity of modeling multiple speakers, we propose Role-Relative Projection, which maps any N-speaker interaction onto a fixed current versus next floor-holder state. Because existing audiovisual datasets contain disruptive editing cuts that break causal tracking, we introduce the Audio-Visual Conversation Corpus, a 31-hour dataset of unedited, single-camera multiparty conversations. Evaluations demonstrate that MuVAP outperforms strong baselines on Shift-Hold and next-speaker prediction tasks across two- and three-speaker settings.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes