CLSep 4, 2025

OleSpeech-IV: A Large-Scale Multispeaker and Multilingual Conversational Speech Dataset with Diverse Topics

Wei Chu, Yuanzhe Dong, Ke Tan, Dong Han, Xavier Menendez-Pidal, Ruchao Fan, Chenfeng Miao, Chanwoo Kim, Bhiksha Raj, Rita Singh

arXiv:2509.04702v12.7h-index: 10Has Code

Originality Synthesis-oriented

AI Analysis

This dataset addresses the need for diverse conversational speech data for researchers in speech processing, though it is incremental as part of a series.

The authors introduced OleSpeech-IV, a large-scale multispeaker and multilingual conversational speech dataset with diverse topics, derived from publicly-available English audio sources and processed with human-sourced and proprietary methods, and they open-sourced a subset for non-commercial research.

OleSpeech-IV dataset is a large-scale multispeaker and multilingual conversational speech dataset with diverse topics. The audio content comes from publicly-available English podcasts, talk shows, teleconferences, and other conversations. Speaker names, turns, and transcripts are human-sourced and refined by a proprietary pipeline, while additional information such as timestamps and confidence scores is derived from the pipeline. The IV denotes its position as Tier IV in the Olewave dataset series. In addition, we have open-sourced a subset, OleSpeech-IV-2025-EN-AR-100, for non-commercial research use.

View on arXiv PDF

Similar