ASAIJul 6

ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions

arXiv:2607.052764.5
Predicted impact top 62% in AS · last 90 daysOriginality Incremental advance
AI Analysis

This work enables controllable generation of speaker embeddings from text descriptions, addressing the need for fine-grained speaker attribute control in speech synthesis systems.

ProPS generates distributions of speaker embeddings conditioned on natural language prompts (e.g., 'a thirties male speaker with an Indian accent') using a mixture density network. It achieves high fidelity in preserving requested attributes like age, gender, and accent, enabling controllable speaker-profile synthesis for TTS/VC.

Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extractors are typically descriptive rather than generative: they map an observed speech segment to an x-vector, which is then used for downstream applications. We introduce ProPS, Prompted Profile Synthesis, a framework for generating distributions of speaker embeddings conditioned on natural language prompts such as "a thirties male speaker with an Indian accent". ProPS converts human-written profile descriptions into sentence embeddings and uses a mixture density network trained on a large-scale dataset to predict a Gaussian mixture model in the x-vector space. The model is trained by maximizing the likelihood that real speaker embeddings match the requested profile, and its generated distributions are evaluated by negative log-likelihood on held-out x-vectors and by attribute classification accuracies on sampled synthetic x-vectors. Experiments show that ProPS produces profile-conditioned distributions and generates x-vectors that preserve requested speaker attributes such as age, gender, accent, and prosodic characteristics. This design enables controllable speaker-profile synthesis for speech generation systems like Text-To-Speech (TTS) or Voice Conversion (VC) while anchoring generated distributions in observed speaker-embedding structure.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes