AI CYJan 30, 2025

Normative Evaluation of Large Language Models with Everyday Moral Dilemmas

arXiv:2501.18081v120.226 citationsh-index: 6FAccT

Originality Synthesis-oriented

AI Analysis

This addresses the problem of assessing LLM moral reasoning for applications like therapy and companionship, though it is incremental by using existing data in a new evaluation framework.

The study evaluated large language models (LLMs) on everyday moral dilemmas from Reddit's 'Am I the Asshole' community, finding that LLMs show distinct moral judgment patterns with moderate to high self-consistency but low inter-model agreement, differing substantially from human evaluations.

The rapid adoption of large language models (LLMs) has spurred extensive research into their encoded moral norms and decision-making processes. Much of this research relies on prompting LLMs with survey-style questions to assess how well models are aligned with certain demographic groups, moral beliefs, or political ideologies. While informative, the adherence of these approaches to relatively superficial constructs tends to oversimplify the complexity and nuance underlying everyday moral dilemmas. We argue that auditing LLMs along more detailed axes of human interaction is of paramount importance to better assess the degree to which they may impact human beliefs and actions. To this end, we evaluate LLMs on complex, everyday moral dilemmas sourced from the "Am I the Asshole" (AITA) community on Reddit, where users seek moral judgments on everyday conflicts from other community members. We prompted seven LLMs to assign blame and provide explanations for over 10,000 AITA moral dilemmas. We then compared the LLMs' judgments and explanations to those of Redditors and to each other, aiming to uncover patterns in their moral reasoning. Our results demonstrate that large language models exhibit distinct patterns of moral judgment, varying substantially from human evaluations on the AITA subreddit. LLMs demonstrate moderate to high self-consistency but low inter-model agreement. Further analysis of model explanations reveals distinct patterns in how models invoke various moral principles. These findings highlight the complexity of implementing consistent moral reasoning in artificial systems and the need for careful evaluation of how different models approach ethical judgment. As LLMs continue to be used in roles requiring ethical decision-making such as therapists and companions, careful evaluation is crucial to mitigate potential biases and limitations.

View on arXiv PDF

Similar