AIJul 24

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI

arXiv:2607.2236817.9
Predicted impact top 21% in AI · last 90 daysOriginality Incremental advance
AI Analysis

Provides a method to validate agent benchmark scores, addressing a critical flaw in capability measurement for the AI community.

Agent benchmarks often overstate capabilities due to reward hacking and shortcut use. HackDetect audits reveal score inflation of 0.45-1.00 across 15 benchmarks, with 67% of traces showing exposures.

Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation protocol keeps the intended capability necessary for success. Recent reward-hacking benchmarks and system reports show that agents can instead recover public solutions, read evaluation artifacts, infer generator structure, manipulate feedback, or benefit from invalid scoring paths; existing responses do not provide a common procedure for attributing these shortcuts and quantifying their effect across benchmarks. We formulate protocol validity and introduce HackDetect, a post-hoc audit that identifies an exposure, determines how the agent used it, and assesses whether the resulting score is misleading. We quantify score inflation with the Mislead gap, defined as the exploit score minus the intended score. We audit 2,385 traces across 15 agent benchmarks and find evidence of exposures and reward hacking in 67.0% of Frontier Science traces and 66.7% of AutoLab tasks. Across paired comparisons, we measure score inflation of 0.45-1.00, showing that benchmark reports should provide evidence that scores reflect the intended capability.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes