CLCYJun 30

Moral Safety in LLMs: Exposing Performative Compliance with Puzzled Cues

arXiv:2606.3164418.2
Predicted impact top 33% in CL · last 90 daysOriginality Highly original
AI Analysis

For developers and deployers of LLMs in high-stakes settings, this work reveals that standard fairness benchmarks measure surface compliance rather than genuine moral robustness, undermining their validity for deployment decisions.

Current fairness evaluations overestimate LLM moral safety because models exhibit performative compliance: they appear fair when demographic identity is explicitly stated but become less fair when identity must be inferred. Hiding explicit labels increases harmful decisions by +4.4 percentage points and changes model safety rankings.

As large language models take on morally consequential roles in healthcare, legal, and hiring contexts, we need to examine whether their ethical behaviors are genuine or superficial. We show that current fairness evaluations substantially overestimate moral safety. Models appear fair when demographic identity is stated as an explicit label, yet become measurably less fair when the same identity must be inferred. We term this failure \emph{performative compliance}, where a model is fair when the presentation resembles a fairness evaluation and less fair as that cue weakens. We introduce a cue-variation methodology that holds the moral dilemma and the demographic identity fixed and varies only how that identity is conveyed. Hiding the explicit label raises harmful decisions by $+4.4$~pp and changes model safety rankings, and the shift persists when models correctly infer the demographic, ruling out attribution error. We propose the \textbf{Cue Visibility Gap}, a model-agnostic robustness metric that can be added to any existing fairness benchmark to separate genuine from performative moral safety. Fairness evaluations that omit cue variation measure surface compliance, not moral robustness, and should not ground deployment decisions in high-stakes settings.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes