SDHCLGAug 1

SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces

arXiv:2608.008033.2
Predicted impact top 87% in SD · last 90 daysOriginality Synthesis-oriented
AI Analysis

This provides a new benchmark dataset for researchers developing wearable silent speech interfaces, enabling open-vocabulary recognition for the first time.

The authors introduce SoniSpeech, the first large-scale open-vocabulary trimodal dataset for wearable silent speech interfaces using acoustic-sensing eyewear, comprising 34 hours and 18,000 utterances. A CTC-based ResNet-34 baseline achieves 26.3% WER on open-vocabulary silent speech recognition, establishing the first benchmark for this task.

Wearable silent speech interfaces (SSIs) are limited to small, closed vocabularies. Approaches achieving larger vocabularies require obtrusive hardware such as facial electrodes. We present SoniSpeech, the first large-scale, open-vocabulary, trimodal dataset for wearable SSI using acoustic-sensing eyewear. It contains 34 hours across 18,000 utterances with three synchronized modalities: ultrasound echo profiles, voiced audio, and frontal video, in both voiced and silent modes. The corpus draws from the SODA dialogue dataset, providing contemporary conversational English with 5,356 unique words and full phoneme coverage. A CTC-based ResNet-34 baseline achieves 26.3% word error rate (WER) on open-vocabulary silent speech recognition, the first benchmark for this task. Dataset is available at https://doi.org/10.7298/xjjr-9m85

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes