Han Zhu, Lingxuan Ye, Wei Kang et al.
This work addresses the challenge of creating a massive multilingual TTS system for broad language coverage, representing a significant advancement rather than an incremental improvement.
Speech recognition, audio signal processing
Han Zhu, Lingxuan Ye, Wei Kang et al.
This work addresses the challenge of creating a massive multilingual TTS system for broad language coverage, representing a significant advancement rather than an incremental improvement.
Ke-Han Lu, Szu-Wei Fu, Chao-Han Huck Yang et al.
This work provides empirical grounding for understanding LLMs in audio research, addressing a gap in knowledge for researchers and practitioners in audio AI.
Kai-Wei Chang, Wei-Chih Chen, En-Pei Hu et al.
This addresses a practical limitation for real-world spoken language systems like voice assistants, where controlling response duration can enhance interaction quality, though it is an incremental improvement.
Jingyu Lu, Yuhan Wang, Fan Zhuo et al.
This work addresses evaluation challenges for spoken dialogue systems, offering a novel benchmark and model to better assess conversational quality, though it is incremental in advancing existing reward modeling approaches.
Zhifeng Kong, Sang-gil Lee, Jaehyeon Kim et al.
For researchers and practitioners in multimodal AI, this work provides a single model that excels in both audio and text tasks without sacrificing text performance, addressing the challenge of unified audio-text intelligence.
Shuiyuan Wang, Zhixian Zhao, Hongfei Yue et al.
Provides a more realistic and rigorous benchmark for emotional intelligence in audio language models, addressing limitations of existing synthetic and single-turn benchmarks.
Yi Su, Jisheng Bai, Qisheng Xu et al.
This is an incremental survey that helps researchers and practitioners in audio-centric AI by summarizing existing technologies and providing references for practical applications.
Guanrou Yang, Tian Tan, Qian Chen et al.
For speech AI researchers, WavCube provides a unified representation that bridges the gap between semantic and acoustic features, enabling a single model for both understanding and generation tasks.
Zhifei Xie, Zihang Liu, Ze An et al.
This work addresses the need for a single model that can handle multiple streaming audio tasks (e.g., voice chatting, ASR) in real time, unifying capabilities that were previously separate.
Chi-Yuan Hsiao, Ke-Han Lu, Yu-Kuan Fu et al.
Addresses the problem of generative collapse in reinforcement learning for speech language models, enabling more natural turn-taking without degrading semantic quality.
Yen-Ting Piao, Jay Chiehen Liao, Wei-Tang Chien et al.
This addresses a gap in understanding cross-modal integration for audio-language models, though it is incremental as it focuses on evaluation rather than new methods.
Xi Wang, Jie Wang, Xingchen Song et al.
For TTS researchers and practitioners, it provides an interpretable diagnostic tool to identify fine-grained acoustic artifacts, addressing the lack of explainable evaluation metrics.
Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar et al.
This work advances open-source audio-language models for researchers and practitioners needing robust understanding of speech, sound, and music, with strong real-world generalization.
Shi Lian, Changtao Li, Bohan Li et al.
This work provides an open-source TTS foundation model with strong generation stability, voice cloning, and emotional expressiveness, advancing the state of the art for multilingual speech synthesis.
Yixuan Zhou, Guoyang Zeng, Xin Liu et al.
This work provides a powerful open-source foundation for multilingual and controllable speech generation, advancing the field by unifying diverse capabilities in a single model.
Marianne de Heer Kloots, Martijn Bentum, Hosein Mohebbi et al.
This work provides insights into the internal learning dynamics of speech models, which is incremental for researchers in computational linguistics and speech processing.
Chih-Kai Yang, Yun-Shao Tsai, Yu-Kai Guo et al.
This addresses a critical bottleneck in auditory AI for applications requiring complex audio comprehension, though it is incremental in improving existing models.
Orantqing, Shengpeng Ji, Junlong Tong et al.
This work addresses the problem of enabling more natural and continuous human-AI interaction for a broad range of users and applications, moving beyond conventional turn-based paradigms.
Ming-Hao Hsu, Xiaohai Tian, Jun Zhang et al.
Identifies and resolves a specific bottleneck in speech LLM reasoning, enabling parity with text LLMs on logical tasks.
Chenshuang Zhang, Kyeong Seon Kim, Chengxin Liu et al.
For researchers developing audio-visual LLMs, this work identifies a fundamental limitation in cross-modality understanding between speech and vision, highlighting the need for speech-grounded video comprehension.