SDAIJun 14

Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech

arXiv:2606.1938115.2
Predicted impact top 13% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For researchers working on code-switching ASR, this method improves synthetic data augmentation by explicitly enforcing language-boundary consistency, yielding significant error reductions.

The paper proposes a code-mixing guided preference-learning framework for synthetic speech generation that improves code-switching fidelity, reducing MER from 12.1%/17.8% to 8.9%/14.2% on SEAME DevMAN/DevSGE when fine-tuning Whisper Large.

Code-switch (CS) Automatic Speech Recognition (ASR) remains challenging due to limited availability of high quality CS text-speech pairs for training. Although synthetic data augmentation via Text-to-speech (TTS) has been explored, existing CS TTS approaches primarily optimise reconstruction fidelity and do not explicitly enforce language-boundary consistency, thereby limiting their effectiveness for CS ASR augmentation. This paper proposes a code-mixing guided preference-learning framework that steers synthetic speech generation toward improved code-switching fidelity using the Code Mixing Index (CMI). Experiments on the SEAME Mandarin-English conversational corpus demonstrate that the proposed method enhances the utility of synthetic data for ASR fine-tuning. Specifically, when fine-tuning Whisper Large, the proposed approach reduces Mixed Error Rate (MER) from 12.1%/17.8% to 8.9%/14.2% on the DevMAN and DevSGE sets, respectively.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes