CLAICYJan 20

XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs

arXiv:2601.14063v12 citationsh-index: 7
Originality Incremental advance
AI Analysis

This addresses the need for better evaluation of cultural competence in LLMs for cross-cultural NLP research, though it is incremental as it builds on existing frameworks.

The authors tackled the problem of evaluating cross-cultural reasoning in large language models (LLMs) by introducing XCR-Bench, a benchmark with 4.9k parallel sentences and 1,098 unique culture-specific items, and found that state-of-the-art LLMs show consistent weaknesses in identifying and adapting cultural elements and encode biases.

Cross-cultural competence in large language models (LLMs) requires the ability to identify Culture-Specific Items (CSIs) and to adapt them appropriately across cultural contexts. Progress in evaluating this capability has been constrained by the scarcity of high-quality CSI-annotated corpora with parallel cross-cultural sentence pairs. To address this limitation, we introduce XCR-Bench, a Cross(X)-Cultural Reasoning Benchmark consisting of 4.9k parallel sentences and 1,098 unique CSIs, spanning three distinct reasoning tasks with corresponding evaluation metrics. Our corpus integrates Newmark's CSI framework with Hall's Triad of Culture, enabling systematic analysis of cultural reasoning beyond surface-level artifacts and into semi-visible and invisible cultural elements such as social norms, beliefs, and values. Our findings show that state-of-the-art LLMs exhibit consistent weaknesses in identifying and adapting CSIs related to social etiquette and cultural reference. Additionally, we find evidence that LLMs encode regional and ethno-religious biases even within a single linguistic setting during cultural adaptation. We release our corpus and code to facilitate future research on cross-cultural NLP.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes