SDASJun 23

Quantifying Dimensional Independence in Speech: An Information-Theoretic Framework for Disentangled Representation Learning

arXiv:2602.205922.2h-index: 42
Predicted impact top 94% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For speech researchers, this provides a principled method to quantify disentanglement in acoustic features, though it is an incremental application of existing MI estimation techniques to a new domain.

The paper introduces an information-theoretic framework to quantify cross-dimension statistical dependence in speech features, finding low cross-dimension mutual information (MI) across six corpora (bounds <0.15 nats) and higher Source-Filter MI (0.47 nats), with attribution analysis revealing source dominance for emotion (80%) and filter dominance for linguistic (60%) and pathological (58%) dimensions.

Speech signals encode emotional, linguistic, and pathological information within a shared acoustic channel; however, disentanglement is typically assessed indirectly through downstream task performance. We introduce an information-theoretic framework to quantify cross-dimension statistical dependence in handcrafted acoustic features by integrating bounded neural mutual information (MI) estimation with non-parametric validation. Across six corpora, cross-dimension MI remains low, with tight estimation bounds ($< 0.15$ nats), indicating weak statistical coupling in the data considered, whereas Source--Filter MI is substantially higher (0.47 nats). Attribution analysis, defined as the proportion of total MI attributable to source versus filter components, reveals source dominance for emotional dimensions (80\%) and filter dominance for linguistic and pathological dimensions (60\% and 58\%, respectively). These findings provide a principled framework for quantifying dimensional independence in speech.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes