AIJun 18

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

arXiv:2606.205328.0
Predicted impact top 78% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in expressive text-to-speech, this work provides the first mechanistic understanding of how natural language style captions control voice characteristics in diffusion-based TTS, enabling better diagnosis of failure modes and improved controllability.

This paper introduces cross-attention attribution for speech diffusion models, adapting the DAAM framework to analyze how style captions influence acoustic output in CapSpeech-TTS. Analysis of 3,600 (caption, transcript) pairs reveals that style tokens have lower temporal variance, correlate with F0 and energy, and peak in early steps and deep layers, with maximal network selectivity at layer 17.

Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this is critical for diagnosing failure modes and improving controllability in expressive TTS. We propose cross-attention attribution for speech diffusion models, adapting the DAAM framework to the speech domain for the first time, and apply it to CapSpeech-TTS. Our method extracts per-token heatmaps across 25 layers and 24 ODE steps. We analyze 3,600 (style caption, text transcript) combinations comprising 120 style captions conditioning the generation of 30 text transcripts each, revealing how caption tokens shape waveforms. Results show: (1) style tokens have lower temporal variance than content/function tokens, confirming global conditioning; (2) style attention correlates with F0 and energy; (3) style conditioning peaks in early steps and deep layers; (4) attention entropy reaches its minimum at layer 17, co-occurring with the style importance peak, indicating maximal network selectivity at the most style-critical stage. This is the first study of how natural language influences cross-attention in speech diffusion models

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes