SDCLASJun 23

ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge

arXiv:2606.2464818.7
Predicted impact top 8% in SD · last 90 daysOriginality Incremental advance
AI Analysis

For researchers evaluating LALMs on fine-grained paralinguistic speech tasks, this benchmark reveals a significant gap between LALM and human judgment, highlighting the need for calibration-aware assessment.

The paper introduces ParaPairAudioBench, a pairwise benchmark of 5,175 audio pairs across five paralinguistic dimensions, and shows that current LALM judges lag behind human judgments by 32%p on average and exhibit severe calibration failures, particularly in Tie cases.

Large Audio-Language Models (LALMs) have been widely used as judge models for the automatic evaluation of generated speech. However, prior approaches predominantly focus on holistic naturalness, leaving fine-grained paralinguistic distinctions underexplored. We introduce ParaPairAudioBench, a pairwise benchmark of 5,175 audio pairs across five paralinguistic dimensions: Style, Rate, Emphasis, Age, and Gender. Our experiments show that current LALM judges still lag behind human judgments by 32%p on average and exhibit severe calibration failures, particularly in Tie cases where the correct decision is to abstain. To further analyze lexical versus acoustic reliance, the benchmark includes both same-transcript and cross-transcript conditions. ParaPairAudioBench enables multi-dimensional, calibration-aware assessment of the reliability of LALM-as-a-Judge for paralinguistic speech evaluation.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes