ASSDJun 20

Using Phonological-Level Wav2Vec2 for Mandarin Automatic Mispronunciation Detection and Diagnosis

arXiv:2606.220223.2
Predicted impact top 96% in AS · last 90 daysOriginality Incremental advance
AI Analysis

Improves diagnostic feedback for L2 Mandarin learners by explicitly separating segmental and tonal errors.

The paper proposes a phonological feature-based MDD framework using Wav2Vec2 CTC that separately models segmental and tonal errors, reducing FAR by 10.1% and DER by 23.6% over a phoneme-only baseline.

Automatic mispronunciation detection and diagnosis (MDD) plays a crucial role in L2 Mandarin pronunciation learning. While end-to-end (E2E) based MDD methods have substantially improved phoneme-level detection accuracy, diagnostic feedback remains limited, as segmental and tonal errors are not explicitly separated. In this paper, we propose a phonological feature-based MDD framework that models both segmental and tonal attributes within a unified Wav2Vec2 CTC architecture. Experimental results show that the proposed method reduces the False Acceptance Rate (FAR) by 10.1% and the Diagnostic Error Rate (DER) by 23.6% compared with the phoneme-only baseline system. By decomposing phonemes into low-level phonological components, the proposed approach enables more detailed and interpretable diagnostic feedback for L2 learners.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes