AICLCYJun 12

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

arXiv:2606.14516v120.3
Predicted impact top 20% in AI · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the fragmentation of AI evaluation results for researchers and practitioners, enabling cross-community analysis and reducing redundancy.

Every Eval Ever introduces a unified schema and community repository for AI evaluation results, standardizing diverse formats and frameworks to enable comparison and reuse. The repository currently includes 22,235 models, 2,273 benchmarks, and 31 evaluation formats.

AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First, results are saved in incompatible formats, scattered across leaderboards, papers, blog posts, evaluation harness logs, and custom repositories. Second, results are created by different evaluation frameworks, which produce divergent scores for nominally identical evaluations and record metadata inconsistently, hindering comparison, cross-community evaluation science, cost reduction, and reuse. We introduce Every Eval Ever, the first shared schema and community-crowdsourced repository for AI evaluation results. The schema standardizes how evaluations are represented in a unified, single JSON document. It is source-agnostic by design, ingesting results from evaluation harnesses and papers alike, and optionally stores per-instance outputs for fine-grained analysis. We contribute: (i) a community-governed metadata schema with a companion instance-level schema, the first standardization effort of its kind; (ii) automatic converters from popular formats, evaluation harnesses, and leaderboards to the unified schema; and (iii) a crowdsourced community database hosted on Hugging Face, currently spanning to date 22,235 models, 2,273 unique benchmarks, and 31 evaluation formats.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes