AIJun 29

Sequential Fairness Auditing with Limited Output Access

arXiv:2606.303383.1
Predicted impact top 95% in AI · last 90 daysOriginality Incremental advance
AI Analysis

It provides a practical statistical framework for independent auditors to evaluate AI fairness under realistic query constraints, addressing a key bottleneck in external AI governance.

This paper formulates fairness auditing as a sequential hypothesis-testing problem under limited model output access, developing a sequential generalized likelihood-ratio framework that reduces query counts by up to 50% for certain fairness metrics compared to fixed-sample methods.

External evaluations are becoming increasingly central to the governance of AI systems. In practice, however, independent auditors often have limited access to deployed models and must rely on query-based interactions. Most existing fairness evaluation methods assume static datasets and fixed-sample statistical tests, making them poorly suited to real-world auditing scenarios in which evidence must be collected sequentially under query constraints. In this work, we formulate fairness auditing as a tolerance-aware sequential hypothesis-testing problem under limited model output access. We develop a sequential generalized likelihood-ratio framework that allows auditors to accumulate evidence from a finite audit pool and stop once sufficient support for compliance or violation has been obtained. The framework is instantiated for decision-based Statistical Parity and Equal Opportunity audits, and extended to score- and logit-based proxy audits when richer observables are available. Our results show that both the fairness metric and the level of model access significantly affect audit efficiency, and that the benefits of richer output information are not uniform across auditing settings. In particular, richer outputs can substantially reduce the number of queries required for some fairness metrics and operating regimes, while offering limited gains in near-threshold cases. This work provides a practical statistical framework for sequential fairness auditing under realistic deployment constraints.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes