CLLGSDASJul 13, 2022

Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS

arXiv:2207.06000v118 citationsh-index: 17
Originality Incremental advance
AI Analysis

This addresses a practical limitation in TTS for users who lack reference speech, enabling more flexible style control.

The paper tackles the problem of controlling emotional style in text-to-speech without requiring target speaker recordings, by proposing a text-based interface and bi-modal style encoder, achieving high-quality expressive speech in unseen styles.

Expressive text-to-speech has shown improved performance in recent years. However, the style control of synthetic speech is often restricted to discrete emotion categories and requires training data recorded by the target speaker in the target style. In many practical situations, users may not have reference speech recorded in target emotion but still be interested in controlling speech style just by typing text description of desired emotional style. In this paper, we propose a text-based interface for emotional style control and cross-speaker style transfer in multi-speaker TTS. We propose the bi-modal style encoder which models the semantic relationship between text description embedding and speech style embedding with a pretrained language model. To further improve cross-speaker style transfer on disjoint, multi-style datasets, we propose the novel style loss. The experimental results show that our model can generate high-quality expressive speech even in unseen style.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes