Ke-Han Lu, Szu-Wei Fu, Chao-Han Huck Yang et al.
This work provides empirical grounding for understanding LLMs in audio research, addressing a gap in knowledge for researchers and practitioners in audio AI.
Audio processing, speech, music
Ke-Han Lu, Szu-Wei Fu, Chao-Han Huck Yang et al.
This work provides empirical grounding for understanding LLMs in audio research, addressing a gap in knowledge for researchers and practitioners in audio AI.
Yitian Gong, Botian Jiang, Yiwei Zhao et al.
This work addresses the need for efficient and controllable text-to-speech systems, though it appears incremental as it builds on existing tokenization and transformer methods.
Junchao Liao, Zhenghao Zhang, Xiangyu Meng et al.
This addresses the challenge of physical coherence in audio-video generation for applications like media production, though it is incremental as it builds on existing methods with a novel kinematic prior.
Beomsoo Kim, Changho Choi, Dohyun Kim et al.
For researchers and developers of speech AI, this work provides a top-performing open-source speech language model and full-duplex system, though it is an incremental improvement over existing architectures.
Xiaofei Wen, Wenjie Jacky Mo, Xingyu Fu et al.
This work addresses the critical problem of audio-visual misalignment in multimodal LLMs for researchers and practitioners building video understanding systems, revealing a systematic failure mode and providing a diagnostic and mitigation approach.
Zhifeng Kong, Sang-gil Lee, Jaehyeon Kim et al.
For researchers and practitioners in multimodal AI, this work provides a single model that excels in both audio and text tasks without sacrificing text performance, addressing the challenge of unified audio-text intelligence.
Dingdong Wang, Shujie Liu, Tianhua Zhang et al.
This work addresses the problem of explainable emotion understanding in speech for applications in multimodal AI, representing a novel approach rather than an incremental improvement.
Tara Bogavelli, Gabrielle Gauthier Melançon, Katrina Stankiewicz et al.
For researchers and developers of voice agents, this benchmark provides a standardized, end-to-end evaluation tool that reveals significant performance gaps and robustness issues not captured by existing benchmarks.
Shuiyuan Wang, Zhixian Zhao, Hongfei Yue et al.
Provides a more realistic and rigorous benchmark for emotional intelligence in audio language models, addressing limitations of existing synthetic and single-turn benchmarks.
Yi Su, Jisheng Bai, Qisheng Xu et al.
This is an incremental survey that helps researchers and practitioners in audio-centric AI by summarizing existing technologies and providing references for practical applications.
Trung Nguyen Quang, Cheng Yi Lewis Won, Minh Duc Pham et al.
Addresses a practical failure mode in multilingual Audio LLMs for code-switching transcription, a known bottleneck in speech recognition.
Santosh Kesiraju, Bolaji Yusuf, Šimon Sedláček et al.
Provides a diagnostic tool for practitioners to understand biases in sentence embeddings without downstream tasks, addressing interpretability in multimodal multilingual models.
Yan Zhou, Qingkai Fang, Yun Hong et al.
This work addresses the challenge of building speech LLMs for multiple languages with limited data, offering a more scalable approach for cross-lingual speech interaction.
Zijian Ling, Pingyi Hu, Xiuyong Gao et al.
This addresses a critical security problem for users of speech-driven LLMs by demonstrating practical, black-box attacks that are perceptually undetectable, though it is incremental in applying known acoustic techniques to a new domain.
Zhifei Xie, Zihang Liu, Ze An et al.
This work addresses the need for a single model that can handle multiple streaming audio tasks (e.g., voice chatting, ASR) in real time, unifying capabilities that were previously separate.
Matan Ben-Yosef, Tavi Halperin, Naomi Ken Korem et al.
This addresses the need for modular and efficient control in audio-visual generation for researchers and practitioners, offering a significant improvement over monolithic or costly methods.
Chi-Yuan Hsiao, Ke-Han Lu, Yu-Kuan Fu et al.
Addresses the problem of generative collapse in reinforcement learning for speech language models, enabling more natural turn-taking without degrading semantic quality.
Yen-Ting Piao, Jay Chiehen Liao, Wei-Tang Chien et al.
This addresses a gap in understanding cross-modal integration for audio-language models, though it is incremental as it focuses on evaluation rather than new methods.
Yuheng Chen, Qingdong He, Teng Hu et al.
This work addresses the underexplored problem of multimodal customization for simultaneous identity preservation in audio-video generation, offering a solution for content creators needing consistent character voices and appearances.
Jielin Qiu, Zixiang Chen, Liangwei Yang et al.
This provides a practical solution for enterprises needing self-hosted, low-latency voice agents, though it is incremental as it builds on existing components rather than introducing new methods.