SDAIJun 19

Imitation Learning for Elder-Facing Speech Synthesis

arXiv:2606.210538.6
Predicted impact top 50% in SD · last 90 daysOriginality Incremental advance
AI Analysis

Addresses the need for age-adapted speech synthesis for older adults, a domain-specific problem with limited prior work.

The paper proposes an imitation learning framework with two-stage on-policy reward learning to improve text-to-speech synthesis for older adults, outperforming baselines in objective and subjective metrics.

Recent advances in text-to-speech (TTS) synthesis have achieved highly natural and expressive speech generation. However, these systems are designed for general adults and overlook older adults' speech comprehension needs due to age-related sensory and cognitive decline. Prior work involves older adults by collecting preference feedback to tune model parameters. However, obtaining sufficient preference data is costly and difficult, as older adults quickly become fatigued during collection. In this paper, we propose a novel imitation learning (IL) framework to learn TTS models from expert demonstrations. We further improve Group Relative Policy Optimization (GRPO) with two-stage on-policy reward learning (OPRL) to mitigate reward hacking under limited supervision from expert demonstration. Experimental results show that GRPO w/ OPRL outperforms GRPO and supervised baselines in objective and subjective metrics. Audio samples are available at https://dongru1.github.io/demo/im-efss

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes