CVJul 15

Audio-Text Cross-Attention with Psycholinguistic Support Features for Ambivalence/Hesitancy Recognition

arXiv:2607.133455.7h-index: 19Has Code
Predicted impact top 70% in CV · last 90 daysOriginality Synthesis-oriented
AI Analysis

This work addresses the specific task of recognizing ambivalence/hesitancy in videos for the ABAW challenge, offering a competitive method that relies only on audio and text.

The authors developed an audio-text system for ambivalence/hesitancy recognition, excluding visual frames. Their ensemble achieved 0.875 average precision and 0.72 macro-F1 on the public development set.

We present an audio-text system for the Ambivalence/Hesitancy Video Recognition Challenge of the 11th ABAW Competition. The method excludes visual frames and represents each video as overlapping 5-second windows aligned with transcript timestamps. Each window combines a 320-dimensional prosodic audio descriptor, a 768-dimensional emotion-oriented RoBERTa embedding, and 74 handcrafted features capturing uncertainty, hedging, and attitudinal conflict. Audio and text are fused via temporal cross-attention, while support features are injected prior to gated multiple-instance learning (MIL) pooling to modulate the window's importance. Predictions from five independently initialized models are averaged. On the labeled public development set, the ensemble achieved an average precision of 0.875 and a macro-F1 of 0.72. Our source code is publicly available at https://github.com/Liga-de-IA-PUCPR/abaw-11-ah-challenge/.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes