MMCLLGJul 19

EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation

arXiv:2607.183363.5h-index: 3INTERSPEECH
Predicted impact top 82% in MM · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in multimodal emotion recognition, this work addresses the overlooked issue of modality-specific uncertainty, improving fusion robustness.

EmoEUS introduces explicit uncertainty supervision for multimodal emotion recognition in conversation, dynamically weighting modalities via learned variance estimates. It outperforms state-of-the-art methods on IEMOCAP and MELD benchmarks.

Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance. However, existing fusion approaches in MERC often ignore modality-specific uncertainty across utterances caused by conflicting cues, varying noise, and missing modality-specific signals. We propose EmoEUS, an explicit uncertainty supervision framework for MERC. EmoEUS performs uncertainty-aware multimodal fusion by dynamically weighting modalities using learned variance estimates. We also introduce an explicitly supervised loss that aligns each utterance's predicted variance with the distance between the utterance's distributional representation and its emotion- and modality-specific cluster center. Experiments on IEMOCAP and MELD show that EmoEUS consistently outperforms state-of-the-art methods.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes