CLAIJul 9

When Synthetic Speech Is All You Have: Better Call GRPO

arXiv:2607.0840935.5h-index: 18
Predicted impact top 2% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For developers of ASR systems in regulated domains where real speech is unavailable, this work shows that reinforcement learning (GRPO) significantly outperforms supervised fine-tuning when adapting to synthetic speech.

In privacy-constrained domains like banking, synthetic speech is a substitute for real recordings but suffers from acoustic mismatch. Using GRPO (a reinforcement learning method) instead of supervised fine-tuning reduces WER by 40% relative (from 36.71% to 22.09%), and combining SFT with GRPO achieves a 45% reduction.

LLM-based ASR adapted to regulated domains such as banking is bottlenecked by privacy: real speech is costly and legally constrained to collect, making synthetic text-to-speech (TTS) an attractive substitute. Yet synthetic speech stays acoustically mismatched with real recordings, and work on this gap has stayed within supervised fine-tuning (SFT). We instead turn to reinforcement learning, and show that Group Relative Policy Optimization (GRPO) extracts far more from the same synthetic speech than SFT. Synthetic-only adaptation of the model with GRPO, a critic-free method rewarding low-WER hypotheses, reduces WER by 40\% relative to SFT (36.71\%$\to$22.09\%), and an SFT-then-GRPO combination pushes this further to 45\%. We trace the gain to behavior rather than representation: GRPO reduces insertion errors by improving stopping calibration and speech-to-text alignment by better anchoring attention to audio, leaving early-layer representations intact. When synthetic speech is the main resource, reinforcement learning should be preferred over supervised fine-tuning.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes