CLAIJun 30

Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG

arXiv:2606.3098923.6
Predicted impact top 10% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For developers and users of LLMs, this work addresses a persistent fairness failure mode with a practical mitigation method that shows cross-model and cross-task transfer.

The paper identifies 'deductive stereotyping' in LLMs, where models apply population-level statistics to individuals, causing biased yet logically coherent inferences. They propose Fair-GCG to discover injection phrases that improve fairness across benchmarks, generalize to larger models, and reduce bias in open-ended generation.

Warning: This paper contains several toxic and offensive statements. While reasoning generally improves fairness in recent large language models (LLMs), failures persist. In this work, we identify a failure mode, deductive stereotyping, in which models apply population-level statistical regularities to individual cases, producing logically coherent yet socially biased inferences. We provide a statistical interpretation of this phenomenon. To steer models toward fairness-aware reasoning, we propose a reasoning-time injection framework. We further introduce Fair-GCG to systematically discover effective injection phrases. Injection phrases discovered by Fair-GCG improve performance across multiple fairness benchmarks, generalize from smaller to larger LLMs, improves reasoning-level fairness, reduces bias in open-ended generation, and transfer to real-world fairness-sensitive tasks.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes