Deriving Benchmarking Datasets from Long-Form Recordings: Challenges and Opportunities
For researchers in child language development and speech processing, this framework enables cross-corpus use and privacy-preserving ML, but the solutions are preliminary and domain-specific.
The paper addresses three challenges in using long-form recordings for child language research: heterogeneous data formats, lack of standardized benchmarks, and privacy constraints. It presents a framework with a standardized dataset collection, replicable benchmarks, and an ethical governance ecosystem, demonstrated via a voice type classification case study.
Long-form recordings (LFRs) of child-centered audio are ecologically valid sources for studying early language development, but three problems limit their use. First, LFR corpora are collected across sites with heterogeneous formats and consent structures, making cross-corpus use non-trivial. Second, without standardized benchmarks, assessing whether tools generalize across languages and conditions is hard. Third, ML workflows rarely respect privacy constraints governing sensitive child speech. This paper presents a framework addressing all three: a standardized collection of 27 child-centered datasets built with open-source tools (S1); a replicable pipeline for four speech-processing benchmarks (S2); and ELSI, a role-based ecosystem embedding ethical governance into the ML workflow (S3). We demonstrate the framework via a voice type classification case study and show the three solutions are mutually dependent.