CLSDJun 10

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction

arXiv:2606.11681v218.3h-index: 1
Predicted impact top 49% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For TTS researchers, UR-BERT enables massively multilingual TTS without G2P resources, but the improvement is incremental over existing methods.

UR-BERT scales multilingual TTS to 495 languages by using Romanized transcription instead of G2P, and improves performance via speech token prediction. It outperforms baselines across many languages and generalizes to unseen ones.

We propose UR-BERT, a Romanized transcription-based text-to-speech (TTS) encoder for massively multilingual TTS systems. Conventional grapheme-to-phoneme (G2P)-based approaches are limited to around 100 languages due to the availability of reliable G2P resources. In contrast, UR-BERT scales to 495 languages by unifying diverse writing systems into a shared Romanization representation. To further enhance phonetic fidelity and text-speech alignment, we introduce a speech token prediction objective during training, which encourages the encoder to learn speech-aware phonetic representations in a data-efficient manner. Experiments show that TTS systems built on UR-BERT consistently outperform recent text encoder baselines across a wide range of languages and resource conditions, and demonstrate strong generalization to unseen languages.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes