EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

arXiv:2606.2065021.2
Predicted impact top 35% in CL · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the need for fine-grained emotional control in text-to-speech for users who require natural language specification of emotions.

EmoInstruct-TTS introduces a dual-path instruction-guided framework for emotional speech synthesis, using Emotion2embed to cover 48 emotional states with fine-grained intensity and an ICE-Flow model to generate emotion representations from free-form instructions. Experiments demonstrate improved emotional controllability and speech naturalness over strong baselines.

Instruction-based controllable speech synthesis enables users to specify emotions through natural language. However, existing approaches often rely on coarse emotion labels and lack explicit modeling of fine-grained intensity. We propose EmoInstruct-TTS, a dual-path instruction-guided framework for emotional speech synthesis. We introduce Emotion2embed, a supervised semantic-acoustic emotion embedding covering 48 emotional states, including fine-grained categories and intensity levels. To infer embeddings from free-form instructions, we design an Instruction-Conditioned Emotion Flow Model (ICE-Flow) that generates acoustically grounded emotion representations. The inferred embeddings are integrated into an LLM-based synthesis pipeline to provide explicit emotional control while preserving semantic planning. Experiments show improved emotional controllability and speech naturalness over strong baselines.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes