SDJun 18

Exploring Pre-training Benefits on Phoneme Addition through Fine-tuning in Speech Synthesis

arXiv:2606.197922.6
Predicted impact top 92% in SD · last 90 daysOriginality Synthesis-oriented
AI Analysis

For TTS researchers, this clarifies the limited transfer benefit of pre-training for phoneme addition, challenging assumptions about fine-tuning efficiency.

The study investigates whether pre-training benefits phoneme addition (learning unseen phonemes) in text-to-speech fine-tuning. Results show pre-training improves naturalness but offers no advantage in phoneme error rate for new phonemes, requiring as much or more data as training from scratch.

Transfer learning is widely used for low-resource text-to-speech. When the target corpus contains phonemes unseen in pre-training, the model must expand its phoneme inventory during fine-tuning; we call the process "phoneme addition." However, it remains unclear whether the pre-trained ability to generate seen phonemes contributes to this process. This study investigates phoneme addition in two settings: (1) a simulation setup using LLM-generated phoneme-controlled corpora that enables investigation without considering confounding factors, and (2) a real-speech cross-lingual transfer setup (English to Japanese) to validate whether the findings hold in practice. Experiments in both settings showed that while fine-tuning achieved higher naturalness than training from scratch, it required as much or more data to achieve comparable PER for new phonemes. These results indicate that pre-training mainly contributes to naturalness improvement, but offers limited benefit for phoneme addition.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes