ASSDJun 13

Dynamic Prosody Prediction in LLM-based TTS for Improving Speaker Similarity

arXiv:2606.152679.4
Predicted impact top 44% in AS · last 90 daysOriginality Incremental advance
AI Analysis

For personalized TTS systems, this work improves speaker similarity by incorporating dynamic prosody prediction, though it is an incremental improvement over existing LLM-based methods.

The paper addresses the lack of style-specific prosody in LLM-based TTS, which limits speaker similarity. By predicting prosody based on previously synthesized speech, they improve speaker similarity across three datasets.

Personalized text-to-speech (TTS) aims to clone the target speaker in the synthesized speech, imitating both the voice and speaking style. Current large language model (LLM)-based TTS methods ignore the style-specific prosodic patterns in generated speech, resulting in deficient style learning and thus limiting speaker similarity in synthesized speech. To this end, we investigate the prosody learning conditioned on the synthesized speech, and propose to predict the prosody of the current syllable based on previously predicted speech. Experimental results obtained on three datasets demonstrated the efficacy of the proposed dynamic prosody prediction method in enhancing the prosody learning capability, thereby improving the speaker similarity of the generated speech. Audio samples are available at https://muzw.github.io/dynapros/.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes