MEAIEMJun 29

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data

arXiv:2606.297845.0
Predicted impact top 68% in ME · last 90 daysOriginality Incremental advance
AI Analysis

For organizations that repeatedly evaluate generative models using noisy crowd-sourced labels, HERO provides a practical way to leverage past data for more accurate and sensitive evaluation.

HERO uses historical evaluation data to reduce bias and variance in generative model performance estimation, improving reliability and sensitivity compared to standard methods. In simulations and real-world benchmarks, it achieves lower mean squared error and better detection of model differences.

Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive to collect and limited in quantity. Organizations thus often turn to collecting vast but noisy "silver" labels from crowdsourced workers or vendor annotators as proxies for gold labels. Because gold remains the evaluation target, naively aggregating noisy silver labels may introduce bias, and estimators built on sparsely observed gold labels may have high variance to resolve the model performance gaps that guide practical decisions. Model evaluation has become an ongoing operational practice rather than a one-time exercise, with evaluation rounds repeating across model versions, releases, and content domains. A natural question is whether the previous historical evaluation data can be used to improve each new round of evaluation. We introduce HERO (History Enhanced RObust model evaluation), a novel framework that uses historical data to suppress bias (improve reliability) and reduce variance (improve sensitivity) in model performance evaluation. HERO calibrates silver labelers' performance learned from historical gold annotations, and stabilizes the resulting estimator by anchoring it to covariate information measured with high precision in the historical data. HERO can be broadly applied across multiple common evaluation tasks, and remains valid when only a subset of historical labelers appears in the current round. We establish conditions under which the bias and variance reductions hold, showcase HERO's performance in simulation studies, and demonstrate its effectiveness on real-world model evaluation benchmarking datasets.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes