MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers
This addresses the need for automated, cost-effective audiobook production without manual prosody configuration or extensive training, though it appears incremental as it builds on existing multimodal and LLM techniques.
The paper tackles the problem of generating expressive audiobooks with consistent speaker prosody by introducing MultiActor-Audiobook, a zero-shot approach that uses multimodal speaker persona generation and LLM-based script instructions, achieving competitive results compared to commercial products in evaluations.
We introduce MultiActor-Audiobook, a zero-shot approach for generating audiobooks that automatically produces consistent, expressive, and speaker-appropriate prosody, including intonation and emotion. Previous audiobook systems have several limitations: they require users to manually configure the speaker's prosody, read each sentence with a monotonic tone compared to voice actors, or rely on costly training. However, our MultiActor-Audiobook addresses these issues by introducing two novel processes: (1) MSP (**Multimodal Speaker Persona Generation**) and (2) LSI (**LLM-based Script Instruction Generation**). With these two processes, MultiActor-Audiobook can generate more emotionally expressive audiobooks with a consistent speaker prosody without additional training. We compare our system with commercial products, through human and MLLM evaluations, achieving competitive results. Furthermore, we demonstrate the effectiveness of MSP and LSI through ablation studies.