CLCYJun 11

Polar: A Benchmark for Evaluating Political Bias in LLMs

arXiv:2606.12922v114.3
Predicted impact top 72% in CL · last 90 daysOriginality Incremental advance
AI Analysis

Provides a reproducible, cross-contextual benchmark for evaluating political bias in LLMs, addressing the need for multilingual and cross-ideological assessment.

The authors introduce Polar, a 4,026-instance multiple-choice benchmark for measuring political bias in LLMs via option-level likelihoods across U.S. and South Korean contexts. Testing 38 LLMs, they find systematic left-progressive bias on U.S. content but more centered patterns on South Korean content, with presentation language alone shifting measured bias.

Political bias in large language models (LLMs) is increasingly significant, but difficult to measure reproducibly across political and linguistic contexts. We introduce Polar, a 4,026-instance multiple-choice benchmark that measures political bias through option-level likelihoods rather than prompt-based generation. Polar covers two ideological axes and eight issue categories derived from the Manifesto Project, and evaluates models in parallel across U.S. and South Korean political contexts. Across 38 LLMs, measured bias varies systematically with political context, issue category, model group, and presentation language. All models lean left-progressive on U.S. political content, but show more centered and mixed patterns on South Korean content. Translation experiments further show that presentation language alone can shift measured bias. These findings highlight the need for multilingual and cross-contextual evaluation of political bias in LLMs.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes