MLLGCOMEJun 17

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning

arXiv:2606.1905711.8
Predicted impact top 16% in ML · last 90 daysOriginality Incremental advance
AI Analysis

For practitioners using LLMs for scalable evaluation, this provides a statistically grounded method to correct biases with minimal human supervision.

The paper addresses systematic biases in LLM-as-a-Judge evaluations, particularly verbosity bias, by formulating the problem as positive-unlabeled learning and proposing a geometric auditing framework using Partial Optimal Transport. The method improves alignment with human preferences and robustness to presentation biases without retraining.

Large Language Models (LLMs) are increasingly used as judges for scalable evaluation, yet such LLM--as--a--Judge systems exhibit systematic biases that are decoupled from semantic quality, most notably verbosity bias. Meanwhile, human supervision is costly and typically selective, yielding reliable positive judgments but leaving most outputs unlabelled and potentially mixed in quality. We formulate LLM evaluation under selective human supervision as a positive--unlabelled learning problem and propose a geometric auditing framework based on Partial Optimal Transport. By aligning a small set of human--verified positives with a reliable subset of unlabelled outputs in a fixed embedding space, our method identifies human--consistent preferences and corrects biased judges without retraining. Experiments demonstrate improved alignment with human preferences, increased robustness to presentation biases, and interpretable confidence estimates, offering a scalable and statistically grounded alternative to existing LLM--as--a--judge pipelines.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes