CLCVJun 18

StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs

arXiv:2606.2052722.3Has Code
Predicted impact top 30% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and developers of multimodal AI systems, this work provides a fine-grained tool to isolate and measure attribute-level social biases, revealing that bias is concentrated in a small set of visual cues.

The authors introduce StylisticBias, a controlled benchmark with ~25K images to measure how specific visual attributes (e.g., age, body type, fashion style) drive social biases in MLLMs. They find that about 15 attributes account for nearly 80% of the total variation in model judgments, with age and body type dominating identity-level effects and fashion style driving attribute-level shifts.

Multimodal large language models (MLLMs) are increasingly deployed in personally and societally consequential settings, yet the visual cues that shape how these models judge people remain poorly understood. Prior work often compares different (groups of) individuals, making it difficult to separate appearance effects from identity differences. We introduce StylisticBias, a controlled benchmark for evaluating attribute-level social bias in MLLMs. We generate 500 photorealistic base faces and create about 50 single-attribute variations per face, producing about 25K images. This design keeps identity fixed and changes one visual attribute at a time. It lets us measure how specific cues shift model judgments. We evaluate six MLLMs across 25 binary social judgment scenarios. We find that age and body type dominate identity-level effects, while fashion style and other visual cues drive the largest attribute-level shifts. We further find that about 15 attributes account for nearly 80\% of the total variation, showing that bias is concentrated in a small set of visual cues. Sensitivity is strongest in judgments that are semantically aligned with appearance, especially socioeconomic and style-related judgments. We release StylisticBias as a benchmark for fine-grained bias evaluation in multimodal models. Code and dataset: https://github.com/timo-cavelius/StylisticBias and https://hf.co/datasets/shaghayegh/stylistic-bias-dataset.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes