CLJul 2

Using embeddings to predict spoken word duration and pitch in Mandarin monosyllabic words

arXiv:2607.0200212.6
Predicted impact top 69% in CL · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers in speech prosody, this demonstrates that embeddings capture durational information, enabling more precise prosody prediction.

The study shows that contextualized embeddings predict spoken word duration for Mandarin monosyllabic words above chance at both type and token levels, and that predicted durations can back-transform normalized f0 contours to the ms time scale, outperforming a permutation baseline.

Time-normalized f0 contours of Mandarin words in conversational speech have been shown to be predictable in part from their contextualized embeddings (CEs). The present study investigates whether CEs also predict spoken word duration for 7470 tokens of Mandarin monosyllabic CV words extracted from a Mandarin corpus of spontaneous speech. We show that CEs indeed are predictive for duration, above chance level, not only at the type level, but also at the level of individual tokens, as indicated by the results obtained with the type-wise and token-wise permutation baselines. We also show that the predicted durations are sufficiently precise to back-transform predicted f0 contours in [0,1] normalized time to contours on the ms time scale. The resulting predicted contours approximate empirical contours and also outperform a permutation baseline.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes