SDJul 22

A Diagnostic Evaluation Framework for AI-Generated Cover Songs Using Music-Theoretic and Acoustic Features

arXiv:2607.196884.3
Predicted impact top 78% in SD · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers and developers of AI music generation systems, this work provides a diagnostic evaluation framework that identifies specific failure modes, but the results are incremental as they confirm the limitations of current automatic metrics.

The paper proposes a five-dimensional diagnostic framework to evaluate AI-generated cover songs, finding that harmonic progression and arrangement have the highest severe-error rates (53% and 47%), while key consistency is better preserved. However, no feature correlation survived multiple-test correction, and an interpretable pilot failed to outperform a fixed baseline, indicating that low-level features cannot replace context-aware musical judgment.

AI-generated covers often fail through local musical errors that a global quality score cannot locate: the vocal contour may remain recognizable while the accompaniment uses the wrong harmonic function, or the output may stay in key while the arrangement remains incomplete. We present a five-dimensional diagnostic framework covering melodic pitch, harmonic progression, key consistency, style consistency, and arrangement/production quality. The benchmark contains 30 covers generated from 5 source songs by 6 systems, with expert severity ratings and 9 symbolic or acoustic features. Harmonic progression and arrangement had the highest severe-error rates (53% and 47%), whereas key consistency was better preserved. Six covers combined acceptable key consistency with severe harmonic errors. Large-leap ratio had a nominal association with melodic ratings (Spearman rho = -0.429, uncorrected p = 0.018), but no feature correlation survived the nine-test multiplicity reference. An interpretable percentile-rule pilot likewise failed to outperform a fixed majority baseline reliably across 16 dimension-level comparisons. The results separate useful diagnostic evidence from dependable automatic scoring: low-level and symbolic summaries can expose particular symptoms, but they do not replace context-aware musical judgment.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes