ASSDJun 30

Beyond Cross-Reconstruction: Probing-Based Disentanglement Evaluation for Acoustic Teleportation Codecs

arXiv:2606.313655.6
Predicted impact top 61% in AS · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in speech disentanglement and audio codecs, this work provides a more reliable evaluation method than cross-reconstruction, highlighting limitations in current training objectives.

The paper extends a probing-based framework to evaluate disentanglement in neural audio codecs for acoustic teleportation, revealing that speaker identity is well-confined but acoustics leak into speech embeddings, with acoustic embeddings estimating room parameters within 0.02 s of supervised baselines.

Some neural audio codecs disentangle speech into latent subspaces encoding content, speaker identity, and acoustics, enabling acoustic teleportation and voice conversion. Existing evaluations rely on cross-reconstruction quality, which cannot reliably detect leakage across partitions. We extend a probing based framework to assess disentanglement by regressing room-acoustic parameters (reverberation time, clarity, and direct-to-reverberant ratio) and classifying speaker identity, using the gap between intended and unintended partitions as the disentanglement measure. Applied to an acoustic teleportation codec, we find speaker identity is largely confined to its partition, while acoustics leak into the speech embeddings due to the training objective. Acoustic embeddings blindly estimate room parameters within 0.02 s of supervised baselines, indicating physically meaningful structure emerges without explicit supervision.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes