AIJun 12

Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability

arXiv:2606.1502913.3
Predicted impact top 48% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For practitioners deploying LLM judges, this method reduces the cost of human annotation needed to verify judge reliability.

Metric Match selects a subset of samples for human annotation to estimate LLM judge reliability, achieving a win-rate of 0.838 against random selection, reducing estimation error by 18.7%, and cutting annotation needs by 32.5%.

LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation. However, the reliability of these judges depends critically on their alignment with human raters -- a property that itself depends on costly human annotations. In this work, we develop a method (Metric Match) for estimating correlation-based reliability metrics of LLM judges from limited annotations. Metric Match selects a subset of samples for human annotation such that the subset matches the population reliability metric with respect to acquired synthetic labels. We empirically show that Metric Match achieves a win-rate of 0.838 against random subset selection across four different correlation metrics and 15 datasets, with an 18.7% decrease in average estimation error and reduces annotation needs by 32.5%. We provide a cost model and highlight a medical case study where our method saves $1,041.67 compared to random selection for expert annotation. Further, we shift our task from reliability estimation to reliability classification of whether a given judge is above a deployment threshold, outperforming random selection with Metric Match. All project code is publicly available, and we additionally provide an installable package for ease of use.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes