CLJun 14

Extending Item Response Theory for Efficient and Meaningful Multilingual Evaluation

arXiv:2606.1564321.6
Predicted impact top 33% in CL · last 90 daysOriginality Highly original
AI Analysis

For researchers evaluating LLMs across many languages, this provides a more efficient and meaningful evaluation framework that handles translation errors and cultural specificity.

Multilingual-IRT extends Item Response Theory to address three issues in multilingual LLM evaluation: linear scaling with languages, translation errors, and conflation of general vs. culture-specific knowledge. It predicts unobserved instances with 11-16% lower binary cross-entropy than baselines, detects translation errors across all languages, and recovers culture-specific items missed by baselines.

Multilingual benchmarks are central to evaluating large language models (LLMs) across languages, but they suffer from three issues: exhaustive evaluation scales linearly with the number of languages, automatic translation introduces errors that are easily missed at scale, and some items conflate general and culture-specific knowledge. We address all three with a unified statistical framework, Multilingual-IRT, which extends Item Response Theory with per-language difficulty deviations, split discriminability separating content from language effects, and per-language ability residuals. Fitting Multilingual-IRT on 25 LLMs across 29 languages of MMLU-Pro-X, we show that its fitted parameters support three practical applications: predicting unobserved (item, LLM, language) instances with 11-16% lower binary cross-entropy than the strongest accuracy-based baseline, surfacing candidate translation errors distributed across all 28 non-English languages, whereas accuracy-based baselines concentrate detections in a few languages, and recovering culture-specific items that accuracy-based baselines miss.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes