SDJun 10

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

arXiv:2606.11611v18.7h-index: 15
Predicted impact top 49% in SD · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the content-fidelity dilemma in speech tokenization for zero-shot TTS, offering a practical solution that balances quality and speed.

SARA proposes a dual-stream VAE that fuses semantic and acoustic representations to overcome the trade-off between content accuracy and audio fidelity in zero-shot TTS, achieving superior reconstruction and natural synthesis with robust performance under accelerated inference.

Zero-shot text-to-speech (TTS) relies on robust speech representations. However, current speech tokenizers face a fundamental trade-off: acoustic codecs preserve high-fidelity audio but lack linguistic constraints, causing content errors during generation, whereas semantic tokens from self-supervised learning (SSL) models ensure precise text alignment but discard some acoustic information. To bridge this gap, we propose SARA, a dual-stream VAE that directly fuses a frozen SSL semantic anchor with a dedicated residual acoustic encoder. This effectively mitigates the dilemma, creating an efficient and compact latent space without relying on complex regularizers. SARA achieves superior reconstruction quality over strong baselines. Furthermore, in downstream zero-shot TTS tasks, it yields highly natural and expressive synthesis quality, and maintains robust generation performance even under accelerated inference, offering a favorable trade-off between synthesis speed and computational cost.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes