An Evaluation Framework for Text-to-Speech Voice Reconstruction
For researchers and practitioners developing TTS systems for speech-impaired users, this framework addresses the lack of sensitive and task-aligned evaluation methods.
The paper proposes an evaluation framework for text-to-speech voice reconstruction that combines subjective Best Worst Scaling with a novel dual-reference distributional measure, demonstrating its reliability across 17 TTS systems and 193 speakers.
Voice reconstruction using Text-to-Speech (TTS) offers a communication method for people with speech disorders, which aims to retain their speaker identity while improving intelligibility. Previous work generally relies on Mean Opinion Score (MOS) to evaluate naturalness and speaker similarity, but this has limited sensitivity and reliability. We propose an evaluation framework with subjective and objective components. Subjectively, we evaluate perceived intelligibility and speaker identity using Best Worst Scaling (BWS) with situational framing. Objectively, we demonstrate that standard measures fail to predict reconstruction success for highly unintelligible speakers, so we introduce a novel dual-reference distributional measure to assess the trade-off between intelligibility and speaker identity. By evaluating the output of 17 zero-shot TTS systems for 193 speakers, we show that our framework provides a reliable and task-aligned approach for assessing voice reconstruction.