Scaling Human and G2P Supervision for Robust Phonetic Transcription
For researchers and practitioners in speech processing, this work identifies the diminishing returns of G2P supervision and demonstrates the effectiveness of ASR pretraining for robust phonetic transcription across diverse speech types.
The paper studies how automatic phonetic transcription performance scales with human and G2P supervision, finding that G2P supervision helps only when fewer than 20-30 hours of human annotation are available, and beyond that, ASR pretraining achieves a 2.3x reduction in weighted phone feature error rate over prior systems.
Expert phonetic annotation is costly, especially for non-standard dialects and atypical speech. A common alternative is using Grapheme-to-Phoneme (G2P) models to auto-generate phonetic labels from text transcripts at scale. We study how automatic phonetic transcription performance scales with human and G2P supervision in English. Using a curated 80-hour benchmark spanning native, non-native and post-stroke speech, we identify a supervision quality threshold: G2P supervision helps only when fewer than 20-30 hours of human annotation are available. Beyond this threshold, it provides no significant benefit and can reduce cross-dialect robustness. What is effective after this threshold is ASR pretraining which we use to achieve a 2.3x reduction in weighted phone feature error rate over prior systems, with strong gains on non-native and aphasic speech. These results suggest that quantity-driven G2P scaling may yield diminishing returns for robust generalization.