AIApr 22

Who Defines Fairness? Target-Based Prompting for Demographic Representation in Generative Models

arXiv:2604.2103625.6h-index: 1

AI Analysis

For users of generative AI, this work provides a transparent, controllable, and accessible method to address demographic biases without requiring model retraining or curated datasets.

The paper proposes a lightweight, inference-time framework that mitigates representational bias in text-to-image models by allowing users to select among multiple fairness specifications, which guide the construction of demographic-specific prompt variants. Across 36 prompts, the method shifts skin-tone outcomes in directions consistent with the declared target and reduces deviation from targets when defined in skin-tone space.

Text-to-image(T2I) models like Stable Diffusion and DALL-E have made generative AI widely accessible, yet recent studies reveal that these systems often replicate societal biases, particularly in how they depict demographic groups across professions. Prompts such as 'doctor' or 'CEO' frequently yield lighter-skinned outputs, while lower-status roles like 'janitor' show more diversity, reinforcing stereotypes. Existing mitigation methods typically require retraining or curated datasets, making them inaccessible to most users. We propose a lightweight, inference-time framework that mitigates representational bias through prompt-level intervention without modifying the underlying model. Instead of assuming a single definition of fairness, our approach allows users to select among multiple fairness specifications-ranging from simple choices such as a uniform distribution to more complex definitions informed by a large language model(LLM) that cites sources and provides confidence estimates. These distributions guide the construction of demographic specific prompt variants in the corresponding proportions, and we evaluate alignment by auditing adherence to the declared target and measuring the resulting skin tone distribution rather than assuming uniformity as 'fairness'. Across 36 prompts spanning 30 occupations and 6 non-occupational contexts, our method shifts observed skin-tone outcomes in directions consistent with the declared target, and reduces deviation from targets when the target is defined directly in skin-tone space(fallback). This work demonstrates how fairness interventions can be made transparent, controllable, and usable at inference time, directly empowering users of generative AI.

View on arXiv PDF

Similar