SDAIJun 16

L-Proto: Language-Aware Episodic Prototypical Training for Multilingual Speaker Verification

arXiv:2606.174165.4
Predicted impact top 73% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For multilingual speaker verification systems, this work addresses the language entanglement problem with a simple training strategy that improves generalization across languages.

Multilingual speaker verification suffers from language-dependent acoustic variability that entangles speaker identity with linguistic characteristics. L-Proto, a language-aware episodic prototypical training strategy, constructs language-consistent episodes to reduce language-driven variation, achieving consistent performance improvements over conventional methods on the TidyVoice Challenge benchmark.

Multilingual speaker verification remains challenging because language-dependent acoustic variability causes speaker identity to become entangled with linguistic characteristics, degrading generalization across languages. In multilingual training, embeddings often encode language cues with speaker identity, causing speakers to form language-specific clusters. We propose L-Proto, a language-aware episodic prototypical training strategy that constructs language-consistent episodes. By sampling speakers from a single language per episode, L-Proto reduces language-driven variation during training and encourages embeddings to focus more directly on speaker identity. Experiments on the TidyVoice Challenge benchmark demonstrate consistent performance improvements over conventional fine-tuning and random episodic sampling across multiple backbone architectures.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes