Measuring Intelligence Beyond Human Scale
This paper addresses the problem of measuring AI intelligence beyond human-level performance, which is crucial for evaluating advanced AI systems, but the proposal is conceptual and lacks empirical validation.
The authors argue that absolute-scale evaluation of intelligence beyond human capability is inherently difficult and propose a new paradigm based on relative measurement, where models generate challenges that separate other systems, yielding an adversarial psychometric rating system that scales with capabilities.
How can we measure intelligence beyond human capability? Human-authored benchmarks saturate, and above human capability, examiners may not know which tasks are both hard and verifiable. We argue that this difficulty is inherent to absolute-scale evaluation and propose a new paradigm based on relative measurement in which models generate public challenges that separate other systems. Aggregating these outcomes yields an adversarial psychometric rating system that can scale with the systems being measured. We describe practical protocols that reduce incentives for private-information attacks, support judge-free adjudication, and naturally scale with agent capabilities. We instantiate the framework across verifiable and open-ended, non-verifiable domains, illustrating how model-generated evaluation can continue to measure systems beyond the human frontier.