SDLGJun 24

Listening Like a Judge: A Music-Aware Framework for Automatic Singing Performance Evaluation

arXiv:2606.264514.6
Predicted impact top 80% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and developers of automated singing evaluation systems, MusicJudge provides a more holistic approach that integrates both lyric and musical dimensions, addressing limitations of prior unimodal methods.

MusicJudge is a framework for automatic singing quality assessment that jointly evaluates lyrical correctness and musical fidelity (pitch and rhythm). It achieves strong agreement with human expert judgments across multiple datasets.

Automatic singing quality assessment (SQA) requires evaluating lyrical correctness and musical fidelity while handling expressive variations. However, existing systems largely rely on either acoustic cues or lyric transcriptions exclusively, limiting holistic performance evaluation. Furthermore, their integration is non-trivial due to challenges in robust singing transcription amid melisma, vibrato, and tempo elasticity. To this end, we propose MusicJudge, a modality-guided framework for automated SQA that performs block-aligned multimodal analysis by coupling lyric correctness with pitch-rhythm fidelity. It detects semantically meaningful lyric blocks using multi-signal matching that integrates semantic embeddings, lexical similarity, and phonetic alignment. To improve singing audio transcription, we introduce Modality-Guided LoRA for ASR fine-tuning. Experiments across datasets demonstrate strong agreement with human expert judgments and validate the generalizability of MusicJudge.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes