Augmenting Dysarthric Speech Severity Assessment with MOS Supervision
For researchers and clinicians needing scalable dysarthria assessment, this work offers a practical augmentation method that reduces reliance on scarce clinical annotations.
The paper addresses the scarcity of clinically annotated dysarthric speech for automatic severity assessment by augmenting training with Mean Opinion Score (MOS) labels from speech synthesis evaluation. Fine-tuning on MOS data improves intelligibility and naturalness prediction, with joint training boosting naturalness, showing that synthesis artifacts and dysarthric speech share perceptual commonalities.
Dysarthria is a speech disorder marked by reduced intelligibility and communicative effectiveness. Automatic utterance-level assessment of dysarthric speech can support scalable speech monitoring and therapy-related analysis. Yet training such systems is bottlenecked by the scarcity of clinically annotated dysarthric speech. This work proposes to augment dysarthric speech assessment using data from speech synthesis evaluations, specifically human-annotated utterances with Mean Opinion Score (MOS) labels from the QualiSpeech corpus. Experiments show that fine-tuning on speech synthesis assessment data consistently improves performance on both intelligibility and naturalness prediction, while joint training yields gains primarily on naturalness. These results suggest that synthesis artifacts and dysarthric speech share perceptual commonalities, and speech synthesis evaluation corpora offer a practical augmentation source that reduces reliance on scarce clinical annotations.