ASSDMay 26, 2020

Contrastive Predictive Coding Supported Factorized Variational Autoencoder for Unsupervised Learning of Disentangled Speech Representations

arXiv:2005.12963v28 citations
AI Analysis

This work addresses unsupervised disentanglement of speech representations, which is incremental as it builds on existing methods like VAEs and contrastive learning.

The paper tackled disentangling style and content in speech signals without supervision, achieving competitive speaker-content separation and improved robustness in phone recognition compared to spectral features.

In this work we address disentanglement of style and content in speech signals. We propose a fully convolutional variational autoencoder employing two encoders: a content encoder and a style encoder. To foster disentanglement, we propose adversarial contrastive predictive coding. This new disentanglement method does neither need parallel data nor any supervision. We show that the proposed technique is capable of separating speaker and content traits into the two different representations and show competitive speaker-content disentanglement performance compared to other unsupervised approaches. We further demonstrate an increased robustness of the content representation against a train-test mismatch compared to spectral features, when used for phone recognition.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes