ASSDJun 15

CraBERT: Efficient Phoneme Encoder Pre-Training via Cascade Fusion of Subword Representations for Text-to-Speech

arXiv:2606.166684.7
Predicted impact top 85% in AS · last 90 daysOriginality Incremental advance
AI Analysis

For TTS researchers, CraBERT reduces pre-training time for phoneme encoders by an order of magnitude while maintaining quality.

CraBERT introduces a cascade-fusion architecture and subword-phoneme alignment to pre-train a phoneme encoder for TTS, achieving comparable MOS scores after ~1 epoch of pre-training versus ~10 epochs for baselines.

This paper introduces CraBERT, a pre-trained phoneme encoder (PPEnc) designed for efficient pre-training in text-to-speech (TTS). CraBERT employs a cascade-fusion architecture and a subword-phoneme alignment algorithm to integrate representations from a pre-trained subword-level BERT into a phoneme-level BERT. This design provides prior word- and sentence-level information, reducing the amount of pre-training required by the phoneme encoder. Subjective listening evaluations show that CraBERT achieves MOS values comparable to existing PPEncs after approximately one epoch of pre-training, whereas the baselines in our comparison are pre-trained for approximately ten epochs. These results demonstrate that CraBERT can efficiently learn representations suitable for improving the perceived naturalness and prosody of synthesized speech.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes