Back to Explore
cs.SDComputer Science

Sound

Audio processing, speech, music

22.3SDMar 18
MOSS-TTS Technical Report

Yitian Gong, Botian Jiang, Yiwei Zhao et al.

This work addresses the need for efficient and controllable text-to-speech systems, though it appears incremental as it builds on existing tokenization and transformer methods.

15.3CLApr 8Code
Raon-Speech Technical Report

Beomsoo Kim, Changho Choi, Dohyun Kim et al.

For researchers and developers of speech AI, this work provides a top-performing open-source speech language model and full-duplex system, though it is an incremental improvement over existing architectures.

17.8CVMay 13Code
When Vision Speaks for Sound

Xiaofei Wen, Wenjie Jacky Mo, Xingyu Fu et al.

This work addresses the critical problem of audio-visual misalignment in multimodal LLMs for researchers and practitioners building video understanding systems, revealing a systematic failure mode and providing a diagnostic and mitigation approach.

45.3CLJul 6
Unified Audio Intelligence Without Regressing on Text Intelligence

Zhifeng Kong, Sang-gil Lee, Jaehyeon Kim et al.

For researchers and practitioners in multimodal AI, this work provides a single model that excels in both audio and text tasks without sacrificing text performance, addressing the challenge of unified audio-text intelligence.

19.6SDMay 13Code208
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

Tara Bogavelli, Gabrielle Gauthier Melançon, Katrina Stankiewicz et al.

For researchers and developers of voice agents, this benchmark provides a standardized, end-to-end evaluation tool that reveals significant performance gaps and robustness issues not captured by existing benchmarks.

20.4SDJun 3
Audio Interaction Model

Zhifei Xie, Zihang Liu, Ze An et al.

This work addresses the need for a single model that can handle multiple streaming audio tasks (e.g., voice chatting, ASR) in real time, unifying capabilities that were previously separate.