ASSDJun 29

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation

arXiv:2606.3094423.6
Predicted impact top 1% in AS · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the challenge of adding speech output to existing S2T LLMs while preserving their original capabilities, which is important for building versatile spoken dialogue systems.

PRIME-Speech enables speech-to-speech generation from a frozen speech-to-text LLM without degrading original S2T performance, achieving accurate spoken responses with low word error rate across speech translation, spoken QA, and multi-turn dialogue.

Strong speech-to-text (S2T) LLMs already provide robust speech perception and text reasoning, but adding speech-to-speech (S2S) output is challenging: fine-tuning the backbone can degrade the original S2T performance, while attaching a downstream talker reintroduces a serial text-to-speech bottleneck. We present PRIME-Speech, a frozen-backbone S2S conversion framework that trains only speech-generation modules. PRIME-Speech synchronizes a causal audio post-decoder with intermediate hidden states of the frozen backbone, so codec tokens are generated from the model's evolving reasoning trajectory rather than from completed text chunks. The post-decoder uses mixed hidden-state, text, and audio-history conditioning, and a training-time packing strategy with turn-level audio KV-cache and position reset stabilizes multi-turn spoken interaction without additional multi-turn S2S training data. Multi-token prediction further reduces the effective codec prediction rate and improves first-audio latency without modifying the reasoning path. Across speech translation, spoken QA, speech understanding, and multi-turn dialogue, PRIME-Speech preserves the S2T behavior of the frozen backbone while producing accurate, low-WER spoken responses.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes