ASSDJun 22

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech

arXiv:2606.2319017.3Has Code
Predicted impact top 12% in AS · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the under-explored application of RL to flow-matching TTS models, offering practical optimizations that yield measurable improvements in objective and subjective metrics.

FlowTTS-GRPO introduces an online reinforcement learning framework for flow-matching-based text-to-speech, achieving improvements in speaker similarity and perceptual quality on CosyVoice 3.0 and F5-TTS, with F5-TTS also showing gains in intelligibility.

Existing Reinforcement Learning (RL) research for Text-to-Speech (TTS) focuses on large language models (LLMs), leaving Flow-Matching (FM) under-explored. We present FlowTTS-GRPO, an online RL framework for FM-based TTS. By converting ordinary differential equation (ODE) trajectories into stochastic differential equation (SDE) paths, our method enables direct fine-tuning of open-source FM models without auxiliary models. We show that a weighted reward combination converges faster than a probabilistic scheme, and identify three practical optimizations: omitting classifier-free guidance (CFG) during training accelerates convergence; synthesizing hard cases improves robustness; and applying RL to the FM component enhances audio-detail metrics. Experiments on CosyVoice 3.0 and F5-TTS demonstrate objective and subjective preference gains in speaker similarity and perceptual quality, with F5-TTS also improving intelligibility.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes