CYMay 7

Picturing Perceptions: An Open-Source Toolkit to Uncover Bias in Humans and Machines

arXiv:2606.136886.2h-index: 4Has Code
Predicted impact top 72% in CY · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and practitioners needing to measure and compare bias across humans and AI systems, this toolkit addresses limitations of traditional bias measures by capturing intersectional identities and enabling AI evaluation.

The paper introduces PictoPercept, an open-source toolkit for measuring bias in humans and AI through visual forced-choice comparisons grounded in population benchmarks. In a study with 283 American adults and GPT-5, they found that humans underestimate Asian American earnings and show ingroup bias only in White males, while GPT-5 exhibits stronger biases, systematically underestimating all female groups.

Bias in human judgment and artificial intelligence systems poses critical challenges across consequential domains like hiring, loans, and criminal justice. However, traditional bias measurement tools face fundamental limitations: they struggle to capture intersectional identities, cannot evaluate AI systems, lack grounding in demographic reality, and remain vulnerable to social desirability effects. We introduce PictoPercept, an open-source toolkit that measures bias through visual forced-choice comparisons grounded in population level benchmarks. Participants view pairs of normed facial photographs and assess who is more likely to have higher earnings, with selections compared against actual U.S. Bureau of Labor Statistics data. We validate PictoPercept with a nationally representative sample of 283 American adults and assess GPT-5, a mainstream generative model, using identical stimuli. Our study reveals three key findings: First, participants dramatically underestimate Asian American earnings despite this group having the highest actual earnings, while overestimating Latino male and White male earnings. Second, ingroup favoritism is not universal as White males show clear ingroup bias, but Asian participants actually underestimate their own group's earnings. Third, GPT-5 exhibits substantially stronger biases than humans, with stark systematic underestimation of all female groups. These findings suggest that PictoPercept enables unified bias assessment across human and AI systems while revealing systematic misperceptions that diverge from demographic reality.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes