ASSDJun 30

How Bilingual Are SSL Speech Models? Cross-Lingual Probing of Articulatory Encoding with Finnish and Russian EMA

arXiv:2606.315277.6
Predicted impact top 48% in AS · last 90 daysOriginality Incremental advance
AI Analysis

This work provides cross-lingual evidence that SSL models encode articulatory information, improving interpretability for speech technology applications.

SSL speech models predict articulatory movements from audio with Pearson r up to 0.68 using only 5 minutes of training data, with multilingual models outperforming monolingual ones, and tongue movements more predictable than lip movements.

SSL speech models capture rich phonetic, prosodic, and acoustic patterns from raw audio, yet how they encode articulatory information across diverse languages remains unclear. Using EMA data from bilingual Finnish-Russian speakers, we evaluate cross-lingual correlations between SSL latent representations and articulatory movements. Models achieve strong prediction performance (Pearson r up to 0.68) even with approximately 5 minutes of training data, with multilingual models outperforming monolingual ones. Intermediate layers encode articulatory features most effectively, and tongue movements are more predictable than lip movements. We also assess the impact of task type (read versus spontaneous speech) and language proficiency, finding higher accuracy for structured tasks and strong generalization across proficiency levels. These results enhance the interpretability of SSL models and show their potential for speech-technology applications.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes