SDJun 18

Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization

arXiv:2606.202188.5
Predicted impact top 53% in SD · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the latency-utility trade-off in streaming voice conversion for real-time applications, offering a novel approach that eliminates algorithmic lookahead.

Zero-VC introduces speaker anonymization as a perturbation mechanism to balance timbre leakage and prosodic utility, enabling strictly causal zero-lookahead streaming voice conversion without degrading quality.

Streaming zero-shot voice conversion struggles to disentangle timbre from linguistic content without degrading utility or inflating latency. Current methods rely on information bottleneck (IB) or speaker perturbation. While IB filters out timbre, it discards prosody, forcing models to explicitly inject features like fundamental frequency. This often requires buffering future frames, creating algorithmic lookahead latency. On the other hand, existing perturbation methods largely overlook the crucial trade-off between timbre leakage and utility preservation. Recognizing this neglected trade-off, we find that the inherent objective of Speaker Anonymization (SA) aligns well with balancing these factors. Thus, we introduce SA as a novel perturbation mechanism to explicitly mitigate timbre leakage while retaining prosodic utility. Crucially, SA's robust representations significantly alleviate the generator's reliance on future context, enabling our strictly causal, zero-lookahead network. Audio samples are available at https://amphionteam.github.io/Zero-VC-demo/.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes