SDAIASMay 19, 2025

MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers

arXiv:2505.13082v14 citationsh-index: 3INTERSPEECH
Originality Incremental advance
AI Analysis

This addresses the need for automated, cost-effective audiobook production without manual prosody configuration or extensive training, though it appears incremental as it builds on existing multimodal and LLM techniques.

The paper tackles the problem of generating expressive audiobooks with consistent speaker prosody by introducing MultiActor-Audiobook, a zero-shot approach that uses multimodal speaker persona generation and LLM-based script instructions, achieving competitive results compared to commercial products in evaluations.

We introduce MultiActor-Audiobook, a zero-shot approach for generating audiobooks that automatically produces consistent, expressive, and speaker-appropriate prosody, including intonation and emotion. Previous audiobook systems have several limitations: they require users to manually configure the speaker's prosody, read each sentence with a monotonic tone compared to voice actors, or rely on costly training. However, our MultiActor-Audiobook addresses these issues by introducing two novel processes: (1) MSP (**Multimodal Speaker Persona Generation**) and (2) LSI (**LLM-based Script Instruction Generation**). With these two processes, MultiActor-Audiobook can generate more emotionally expressive audiobooks with a consistent speaker prosody without additional training. We compare our system with commercial products, through human and MLLM evaluations, achieving competitive results. Furthermore, we demonstrate the effectiveness of MSP and LSI through ablation studies.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes