Back to Explore
eess.ASElectrical Engineering

Audio & Speech Processing

Speech recognition, audio signal processing

11.6CLMar 23Code
TiCo: Time-Controllable Training for Spoken Dialogue Models

Kai-Wei Chang, Wei-Chih Chen, En-Pei Hu et al.

This addresses a practical limitation for real-world spoken language systems like voice assistants, where controlling response duration can enhance interaction quality, though it is an incremental improvement.

45.3CLJul 6
Unified Audio Intelligence Without Regressing on Text Intelligence

Zhifeng Kong, Sang-gil Lee, Jaehyeon Kim et al.

For researchers and practitioners in multimodal AI, this work provides a single model that excels in both audio and text tasks without sacrificing text performance, addressing the challenge of unified audio-text intelligence.

20.4SDJun 3
Audio Interaction Model

Zhifei Xie, Zihang Liu, Ze An et al.

This work addresses the need for a single model that can handle multiple streaming audio tasks (e.g., voice chatting, ASR) in real time, unifying capabilities that were previously separate.

13.3SDJun 5Code
dots.tts Technical Report

Shi Lian, Changtao Li, Bohan Li et al.

This work provides an open-source TTS foundation model with strong generation stability, voice cloning, and emotional expressiveness, advancing the state of the art for multilingual speech synthesis.

26.5SDJun 5Code
VoxCPM2 Technical Report

Yixuan Zhou, Guoyang Zeng, Xin Liu et al.

This work provides a powerful open-source foundation for multilingual and controllable speech generation, advancing the field by unifying diverse capabilities in a single model.

35.4ASSep 8Code159
Omni Interaction Agent Technical Report

Orantqing, Shengpeng Ji, Junlong Tong et al.

This work addresses the problem of enabling more natural and continuous human-AI interaction for a broad range of users and applications, moving beyond conventional turn-based paradigms.