AS CLJul 7, 2024

Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu, Chen Yang, Jiaqi Li, Peiyang Shi, Yuancheng Wang, Kai Chen

arXiv:2407.05361v332.7273 citationsh-index: 12

Originality Synthesis-oriented

AI Analysis

This addresses the problem of limited training data for speech generation models, particularly for researchers and developers in AI and speech technology, though it is incremental as it builds on existing data collection efforts.

The authors tackled the scarcity of large, diverse, and spontaneous speech datasets for speech generation by introducing Emilia, a dataset with over 101k hours of speech across six languages, and Emilia-Pipe, an open-source preprocessing pipeline, which together enable more natural and spontaneous speech generation.

Recent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and spontaneous speech datasets. In response, we introduce Emilia, the first large-scale, multilingual, and diverse speech generation dataset. Emilia starts with over 101k hours of speech across six languages, covering a wide range of speaking styles to enable more natural and spontaneous speech generation. To facilitate the scale-up of Emilia, we also present Emilia-Pipe, the first open-source preprocessing pipeline designed to efficiently transform raw, in-the-wild speech data into high-quality training data with speech annotations. Experimental results demonstrate the effectiveness of both Emilia and Emilia-Pipe. Demos are available at: https://emilia-dataset.github.io/Emilia-Demo-Page/.

View on arXiv PDF

Similar