SDJun 12

Explainable and Trustworthy Speech Emotion Recognition Using Confidence Score and Reinforcement Learning Rectified Speech Emotion Descriptors

arXiv:2606.14086v18.1
Predicted impact top 55% in SD · last 90 daysOriginality Synthesis-oriented
AI Analysis

This work addresses the challenge of reliable SED labels for explainable SER, offering a method to improve post-training performance on existing benchmarks.

The paper tackles explainable and trustworthy speech emotion recognition (SER) by proposing a confidence score and reinforcement learning-based method to rectify automatically annotated speech emotion descriptors. The best system achieves absolute SER gains of 2.9% on IEMOCAP and 3.3% on MELD over baselines.

Explainable and trustworthy speech emotion recognition (SER) remains a challenging task to date, largely due to the scarcity of SER data with reliable speech emotion descriptor (SED) labels, such as prosodic features and speaker traits. This paper presents a confidence score and reinforcement learning (RL) based on-the-fly SED rectification approach for post-training SER systems on automatically annotated SED labels. Experiments on IEMOCAP and MELD suggest that explainable SER systems incorporating the proposed confidence score and RL-based SED rectification approach consistently outperform baselines without data selection or SED rectification. The best performing system, which integrates both components, surpasses the baseline without data selection and SED rectification, achieving SER gains of 2.9% and 3.3% absolute (3.7% and 5.4% relative) on IEMOCAP and MELD benchmarks, respectively.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes