CVAIJul 16

Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach

arXiv:2607.1470210.4h-index: 10
Predicted impact top 38% in CV · last 90 daysOriginality Synthesis-oriented
AI Analysis

This work provides a computationally efficient method for ambivalence recognition in affective computing, but the improvement is incremental over text-only baselines.

The authors tackle automatic recognition of ambivalence and hesitancy using a text-centered multimodal fusion approach, achieving a Macro F1-score of 78.24% on the Private Test subset, outperforming the text-only model by 4.03%.

Automatic recognition of ambivalence and hesitancy is challenging because these states may be expressed through inconsistent linguistic, acoustic, facial, and contextual patterns, while top-performing systems often rely on computationally expensive ensembles. We present a single text-centered multimodal approach for video-level ambivalence and hesitancy recognition for the 11th Affective & Behavior Analysis in-the-Wild (ABAW) Challenge. The proposed approach combines linguistic, acoustic, facial, and scene features using text-centered multimodal fusion model. Text Residual Fusion treats text as the anchor modality and applies gated residual adjustments based on the other modalities. Experiments on the Behavioural Ambivalence/Hesitancy (BAH) corpus confirm that text is the strongest unimodal modality. The Text Residual Fusion model achieves an average Macro F1-score (MF1) of 75.14% across the Development and Public Test subsets. On the Private Test subset, it reaches an MF1 of 78.24%, outperforming the text model by 4.03%. These results demonstrate that complementary multimodal information can improve recognition performance without requiring a large model ensemble.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes