CLJun 16

Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement

arXiv:2606.1750622.6Has Code
Predicted impact top 28% in CL · last 90 daysOriginality Highly original
AI Analysis

For NLP researchers and practitioners, this work highlights a previously unmeasured form of bias in LLMs used as judges, pointing to the need for more theoretically grounded bias evaluation methods.

The paper introduces a novel method to evaluate second-order bias in LLMs—bias in how models judge biased content—using a reasoning task grounded in epistemic entitlement. Results show that this task evades safety guardrails, reveals systematic bias variations across target groups, and reflects implicit social maps.

Evaluations of social bias in LLMs largely focus on whether models generate or imply biased content. However, as LLMs are increasingly used as judges of bias, they may exhibit social biases in subtler ways in how they evaluate biased content, which current methods do not systematically capture. We call this second-order bias: social bias in an LLM's judgment about social bias, which we evaluate through a novel, philosophically grounded reasoning task. Drawing on entitlement epistemology, we conceptualize bias as misplaced foundational knowledge that shapes an agent's rational inquiry, and derive a logical reasoning task for LLMs to judge to whom a biased text is acceptable or non-acceptable. We develop two simple metrics to measure how biased LLM judges are in inferring demographics for acceptability without sufficient support, and how these inferences vary across groups targeted by biased texts. Evaluating open and closed models, we find that our task evades safety guardrails by surfacing bias in model judgment. It varies systematically across target groups, reflects implicit social maps, and shows how models are still triggered by demographic labels. Our work points to the need for LLM bias evaluation in judgment tasks and broadly, for more theoretically grounded approaches to bias evaluation in NLP. We release our code and model responses at https://github.com/uofthcdslab/second-order-bias.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes