SDAIJun 23

ZONOS2 Technical Report

arXiv:2606.2432011.6
Predicted impact top 29% in SD · last 90 daysOriginality Incremental advance
AI Analysis

This work advances text-to-speech for researchers and developers by providing a high-quality, open-source model with improved efficiency and fidelity.

ZONOS2 8B, a TTS model with 8B total parameters (900M active) using a novel MoE backbone, achieves state-of-the-art naturalness, prosody, and voice cloning fidelity, trained on over 6M hours of data, and performs competitively on quality, speaker similarity, WER, and a new TTS benchmark while maintaining good streaming latency.

We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve upon Zonos-v0.1 across scale, data, and training recipe. We scale the model from 1.6B to 8B total parameters (900M active) with a novel mixture-of-experts (MoE) backbone, improving inference latency and throughput. We expand our training corpus from 200K to over 6M hours using a new data processing pipeline, and we simplify our post-training and conditioning recipes to improve naturalness and voice cloning fidelity. We evaluate ZONOS2 8B on quality, speaker similarity, WER, and ZTTS1-Eval, our novel TTS benchmark, where it performs competitively with state-of-the-art systems while maintaining good streaming latency. We release our model weights and example inference code under an Apache 2.0 license on GitHub and Hugging Face.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes