AS CL SDAug 20, 2020

Efficient neural speech synthesis for low-resource languages through multilingual modeling

Marcel de Korte, Jaebok Kim, Esther Klabbers

arXiv:2008.09659v18.023 citations

Originality Incremental advance

AI Analysis

This addresses the challenge of producing high-quality synthetic speech for low-resource languages where abundant data is unavailable, though it appears incremental as it builds on existing multilingual and multi-speaker techniques.

The paper tackled the problem of high data requirements for neural TTS in low-resource languages by investigating multilingual multi-speaker modeling as an alternative to monolingual approaches, finding that it can increase naturalness and achieve comparable results to monolingual models.

Recent advances in neural TTS have led to models that can produce high-quality synthetic speech. However, these models typically require large amounts of training data, which can make it costly to produce a new voice with the desired quality. Although multi-speaker modeling can reduce the data requirements necessary for a new voice, this approach is usually not viable for many low-resource languages for which abundant multi-speaker data is not available. In this paper, we therefore investigated to what extent multilingual multi-speaker modeling can be an alternative to monolingual multi-speaker modeling, and explored how data from foreign languages may best be combined with low-resource language data. We found that multilingual modeling can increase the naturalness of low-resource language speech, showed that multilingual models can produce speech with a naturalness comparable to monolingual multi-speaker models, and saw that the target language naturalness was affected by the strategy used to add foreign language data.

View on arXiv PDF

Similar