SDJun 16

DeSRPA: Decoupled Speech Role-Playing Agent via Inference-Time Intervention

arXiv:2606.1766919.8
Predicted impact top 7% in SD · last 90 daysOriginality Highly original
AI Analysis

This work addresses the problem of poor generalization and reasoning degradation in speech role-playing agents, offering a scalable solution for creating immersive characters without role-specific training data.

DeSRPA introduces a training-free, inference-time intervention framework for speech role-playing that decouples cognitive reasoning from paralinguistic expression, outperforming end-to-end fine-tuned models in personality and emotional consistency while achieving high speech naturalness comparable to GPT-4o Audio.

While Large Language Models (LLMs) have revolutionized text-based role-playing, creating immersive Speech Role-Playing Agents (SRPAs) requires a seamless bridge between cognitive reasoning and paralinguistic nuances. Current SRPAs primarily rely on end-to-end (E2E) fine-tuning. However, this paradigm suffers from poor generalization to unseen characters due to its reliance on role-specific data, while imposing a "modality alignment tax" that degrades intrinsic LLM reasoning capabilities. We propose DeSRPA, an agentic framework for character role play via inference-time intervention on frozen backbones. DeSRPA employs a dual-level control vector mechanism, Internal Cognitive Steering and External Expressive Rendering, to synchronize "mind" and "voice". Experiments on SpeechRole and OmniCharacter benchmarks demonstrate that DeSRPA significantly outperforms E2E baselines in personality and emotional consistency. It achieves high speech naturalness, narrowing the gap with proprietary models like GPT-4o Audio, while remaining a scalable and training-free paradigm.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes