CLSDJul 23

An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations

arXiv:2607.2142414.5
Predicted impact top 56% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in automated audio captioning, this provides a more reliable evaluation method for structured outputs, addressing limitations of existing flat-text metrics.

The paper tackles the challenge of evaluating structured audio captions, proposing a multi-axis framework combining LLM judges and deterministic metrics. The framework successfully distinguishes meaning-preserving paraphrases from genuine corruptions, validated via controlled perturbation testing.

Recent advancements in automated audio captioning (AAC) have shifted from monolithic sentence generation toward structured formats that explicitly disentangle distinct acoustic and semantic properties. However, evaluating this heterogeneous data remains a significant challenge. Existing caption metrics focus on flat textual outputs and fail to reliably assess multimodal attributes. To bridge this gap, we propose a multi-axis evaluation framework tailored for structured audio descriptions. Building on the AudioCards dataset, we evaluate outputs across five orthogonal axes: tag-sets, descriptions, logical reasoning, numeric measurements, and spectral profiles. Our approach combines Large Language Model (LLM) judges to capture semantic nuance with deterministic computational metrics to precisely measure acoustic deviations. To rigorously validate the reliability of this framework, we introduce a controlled perturbation testing protocol that injects typed, graded errors into groundtruth annotations. Our results demonstrate that this framework successfully distinguishes meaning-preserving paraphrases from genuine semantic and acoustic corruptions.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes