CLAIJun 19

LLM-Based Multi-Reference Evaluation for Efficient and Robust Assessment of Phrase Break Annotations

arXiv:2606.2109816.8
Predicted impact top 56% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For speech researchers and engineers, this provides a scalable and robust evaluation method for prosodic phrasing that accounts for multiple valid annotations, addressing a key limitation of existing single-reference approaches.

The paper proposes LLM-based Multi-Reference Evaluation (LMRE) for phrase break annotations, which generates multiple valid phrasings from minimal demonstrations. On a Korean testbed of 1,356 annotations, LMRE shows stronger alignment with human judgment than single-reference evaluation in both acceptance behavior and score correlation.

Reliable evaluation of phrase break annotations is crucial, as subtle variations in prosodic boundaries directly affect the clarity and naturalness of speech. However, existing approaches exhibit major limitations: single-reference evaluation assumes a unique gold phrasing for an utterance despite multiple valid phrasings, while human judgment, though flexible, is labor-intensive and unscalable. To address these, we propose LLM-based Multi-Reference Evaluation (LMRE) for phrase break annotations that models the one-to-many nature of prosodic phrasing and generates multiple valid phrasings from minimal demonstrations. On a Korean testbed of 1,356 annotations covering five strategies, LMRE shows stronger alignment with human judgment than single-reference evaluation in both acceptance behavior and score correlation. Our findings demonstrate that LMRE effectively achieves both scalability and multi-reference support, highlighting the potential of LLMs for evaluation in the speech domain.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes