SDASJun 15

Interpretable Audio Editing Evaluation via Chain-of-Thought Difference-Commonality Reasoning with Multimodal LLMs

arXiv:2509.1697514.31 citationsh-index: 10Has Code
Predicted impact top 16% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in audio quality assessment, this work provides an interpretable and scalable alternative to subjective listening tests.

The paper proposes the first natural language-based automated evaluation framework for audio editing, using Qwen2-Audio with caption-based fine-tuning and Chain-of-Thought prompting. It outperforms existing baselines in aligning with human judgments.

Automatic mean opinion score (MOS) prediction serves as a principled alternative to both subjective listening tests and objective metrics, providing scalable and consistent audio evaluation. Inspired by the LLM-as-Judge paradigm, recent multimodal large language models offer strong perceptual modeling and reasoning capabilities, enabling audio quality assessment. In this work, we address the challenging problem of audio editing evaluation and propose the first natural language-based automated evaluation framework built upon Qwen2-Audio. Two caption-based fine-tuning tasks are introduced to enhance multi-audio understanding, together with a designed Chain-of-Thought prompting strategy to encourage structured, step-by-step reasoning. Experiments show that our framework produces interpretable and logically consistent text-based evaluations, aligning closely with human judgments while outperforming existing baselines. The code and demo are available at https://github.com/NKU-HLT/Eval_Reasoning.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes