Right or Wrong, Models Comply: Directional Blindness in LLM Moral Judgment
For AI alignment researchers, this reveals a distinct failure mode where models comply with misleading moral nudges as readily as helpful ones, suggesting alignment should target directionally calibrated updating.
The paper identifies that LLMs exhibit directional blindness in moral judgments, following helpful and harmful nudges at nearly identical rates (asymmetry score 1.04), unlike factual judgments where they follow helpful nudges more (asymmetry score 1.58). This failure mode persists across models and prompting strategies.
As language models take integrated roles across many domains, the response of LLMs to user pushback becomes a critical alignment property. Yet many existing evaluations treat compliance as unidirectional, measuring whether models resist pressure but not whether they resist it selectively. We introduce Compliance Asymmetry (A = BCR/HCR), a bidirectional diagnostic that compares beneficial output change under helpful nudges with harmful change under misleading nudges. Across 9 models and 972,000 nudge-condition responses, we find that this selectivity differs in factual and moral judgments: models follow helpful nudges more than harmful ones on factual questions (A = 1.58), but follow both directions at nearly identical rates on moral questions (A = 1.04). This phenomenon persists across model families, capability levels, and nudging types. Interestingly, we also find that chain-of-thought prompting amplifies helpful and harmful compliance together, while identity-based prompting suppresses both by nearly identical margins. These results identify direction-blind moral compliance as a distinct failure mode in current LLMs and suggest that alignment should target directionally calibrated updating rather than lower compliance alone.