CLLGSDJun 14

Scaling Human and G2P Supervision for Robust Phonetic Transcription

arXiv:2606.1601916.3
Predicted impact top 60% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and practitioners in speech processing, this work identifies the diminishing returns of G2P supervision and demonstrates the effectiveness of ASR pretraining for robust phonetic transcription across diverse speech types.

The paper studies how automatic phonetic transcription performance scales with human and G2P supervision, finding that G2P supervision helps only when fewer than 20-30 hours of human annotation are available, and beyond that, ASR pretraining achieves a 2.3x reduction in weighted phone feature error rate over prior systems.

Expert phonetic annotation is costly, especially for non-standard dialects and atypical speech. A common alternative is using Grapheme-to-Phoneme (G2P) models to auto-generate phonetic labels from text transcripts at scale. We study how automatic phonetic transcription performance scales with human and G2P supervision in English. Using a curated 80-hour benchmark spanning native, non-native and post-stroke speech, we identify a supervision quality threshold: G2P supervision helps only when fewer than 20-30 hours of human annotation are available. Beyond this threshold, it provides no significant benefit and can reduce cross-dialect robustness. What is effective after this threshold is ASR pretraining which we use to achieve a 2.3x reduction in weighted phone feature error rate over prior systems, with strong gains on non-native and aphasic speech. These results suggest that quantity-driven G2P scaling may yield diminishing returns for robust generalization.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes