SDAICLLGASApr 23, 2021

Deep Learning Based Assessment of Synthetic Speech Naturalness

arXiv:2104.11673v121.882 citationsHas Code
Originality Synthesis-oriented
AI Analysis

This provides a tool for evaluating Text-To-Speech or Voice Conversion systems, but it is incremental as it builds on existing methods for speech quality estimation.

The paper tackles the problem of objectively assessing synthetic speech naturalness by developing a language-independent model based on a CNN-LSTM network, achieving improved reliability through transfer learning from speech quality prediction models trained on POLQA scores.

In this paper, we present a new objective prediction model for synthetic speech naturalness. It can be used to evaluate Text-To-Speech or Voice Conversion systems and works language independently. The model is trained end-to-end and based on a CNN-LSTM network that previously showed to give good results for speech quality estimation. We trained and tested the model on 16 different datasets, such as from the Blizzard Challenge and the Voice Conversion Challenge. Further, we show that the reliability of deep learning-based naturalness prediction can be improved by transfer learning from speech quality prediction models that are trained on objective POLQA scores. The proposed model is made publicly available and can, for example, be used to evaluate different TTS system configurations.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes