NVMOS: Non-Verbal Vocalization Quality Assessment in Speech
For researchers and developers of text-to-speech systems, this work provides a specialized evaluation method for non-verbal vocalizations, filling a gap in existing speech quality assessment.
The paper addresses the underexplored problem of perceptual quality assessment for non-verbal vocalizations (NVs) in speech. They constructed an NV-MOS dataset and proposed NVMOS, the first model to predict NV quality, achieving expert-level agreement with human ratings.
Non-verbal vocalizations (NVs), such as laughter, sighs, and coughs, are important acoustic cues for emotion and intent. Existing speech quality assessment methods typically focus on overall naturalness, while non-verbal TTS evaluations mainly examine whether a target NV appears with the correct type and position. However, the perceptual quality of NV events themselves remains underexplored. To address this gap, we construct an NV-MOS dataset containing outputs from multiple NV-TTS systems and naturally occurring NV samples, with ratings collected from three acoustic experts on a perceptual quality scale. We further analyze audio-capable multimodal large language models such as Gemini and find clear inconsistencies between their scores and expert ratings. These results suggest that general-purpose multimodal models cannot reliably replace human judgments for NV quality assessment. We then propose NVMOS, to our knowledge the first model that can reliably predict the perceptual quality of NV events in speech. Experimental results show that, with a local NV-event focusing module, NVMOS reaches expert-level or stronger agreement with human MOS.