LISE : Listenable Interpretable Speaker Embeddings
Provides a label-free method to interpret speaker embeddings for researchers and practitioners in speaker verification, though the interpretability gains are incremental.
LISE decomposes pretrained speaker embeddings into interpretable components without requiring attribute labels, preserving ASV performance (negligible EER degradation) and achieving 83.9% accuracy in human listening tests for speaker discrimination.
Deep neural network-based automatic speaker verification (ASV) systems achieve impressive performance but their embedding representations remain opaque, lacking a structured and perceptually verifiable explanation of the vocal characteristics they encode. Existing approaches either require annotation of speaker attributes or introduce alternative representations whose interpretability is unvalidated with listeners. We propose Listenable Interpretable Speaker Embeddings (LISE), a label-free framework that decomposes pretrained speaker embeddings into a small set of components. This decomposition yields a structured representation that supports the analysis of what information has been encoded by speaker embeddings. LISE preserves ASV performance with negligible EER degradation on x-vector and ECAPA-TDNN. Crucially, the interpretability of these components for human listeners is demonstrated through listening experiments, where participants distinguished speakers with 83.9% accuracy.