Phoneme-Level Mispronunciation Screening in Polish-Speaking Children with an Explainable Assistant
Provides an explainable, low-resource screening tool for early speech error detection in Polish-speaking children, addressing limited access to specialists.
The paper presents a lightweight screening pipeline for Polish-speaking children to detect sibilant substitutions, using a wav2vec2-based recognizer achieving 88.7% exact sequence match on a test set of 10 children (559 utterances). The screening proxy yields 72.9% precision, 61.4% recall, and F1=0.67 with a 2.7% false-alarm rate.
Early identification of speech sound errors in children is often limited by access to specialists, motivating lightweight screening tools that can operate outside the clinic. We present a screening pipeline for Polish-speaking children focused on sibilant substitutions, coupling a wav2vec2-based CTC token recognizer with alignment-based error typing and a template-grounded caregiver assistant for screening, not diagnosis. On a held-out test set of 10 unseen children comprising 559 utterances, the recognizer achieves 88.7 percent exact sequence match. As a conservative screening proxy, we flag a mismatch when the system emits substitution-evidence bracketed tokens at the target segment, yielding 72.9 percent precision, 61.4 percent recall, F1 = 0.67, and a 2.7 percent false-alarm rate on target-correct items. We describe the assistant's safety boundaries and outline a clinician-in-the-loop validation plan for future deployment.