ZONOS2 Technical Report
This work advances text-to-speech for researchers and developers by providing a high-quality, open-source model with improved efficiency and fidelity.
ZONOS2 8B, a TTS model with 8B total parameters (900M active) using a novel MoE backbone, achieves state-of-the-art naturalness, prosody, and voice cloning fidelity, trained on over 6M hours of data, and performs competitively on quality, speaker similarity, WER, and a new TTS benchmark while maintaining good streaming latency.
We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve upon Zonos-v0.1 across scale, data, and training recipe. We scale the model from 1.6B to 8B total parameters (900M active) with a novel mixture-of-experts (MoE) backbone, improving inference latency and throughput. We expand our training corpus from 200K to over 6M hours using a new data processing pipeline, and we simplify our post-training and conditioning recipes to improve naturalness and voice cloning fidelity. We evaluate ZONOS2 8B on quality, speaker similarity, WER, and ZTTS1-Eval, our novel TTS benchmark, where it performs competitively with state-of-the-art systems while maintaining good streaming latency. We release our model weights and example inference code under an Apache 2.0 license on GitHub and Hugging Face.