CL AINov 4, 2025

POLIS-Bench: Towards Multi-Dimensional Evaluation of LLMs for Bilingual Policy Tasks in Governmental Scenarios

Tingyue Yang, Junchi Yao, Yuhui Guo, Chang Liu

arXiv:2511.04705v1h-index: 1Has Code

Originality Incremental advance

AI Analysis

This addresses the need for better evaluation of LLMs in governmental policy tasks, though it appears incremental as it builds on existing benchmarking approaches.

The authors introduced POLIS-Bench, a new evaluation suite for LLMs in governmental bilingual policy scenarios, and found that reasoning models performed best with superior cross-task stability and accuracy, while their fine-tuned lightweight model achieved parity with or surpassed proprietary baselines at reduced cost.

We introduce POLIS-Bench, the first rigorous, systematic evaluation suite designed for LLMs operating in governmental bilingual policy scenarios. Compared to existing benchmarks, POLIS-Bench introduces three major advancements. (i) Up-to-date Bilingual Corpus: We construct an extensive, up-to-date policy corpus that significantly scales the effective assessment sample size, ensuring relevance to current governance practice. (ii) Scenario-Grounded Task Design: We distill three specialized, scenario-grounded tasks -- Clause Retrieval & Interpretation, Solution Generation, and the Compliance Judgmen--to comprehensively probe model understanding and application. (iii) Dual-Metric Evaluation Framework: We establish a novel dual-metric evaluation framework combining semantic similarity with accuracy rate to precisely measure both content alignment and task requirement adherence. A large-scale evaluation of over 10 state-of-the-art LLMs on POLIS-Bench reveals a clear performance hierarchy where reasoning models maintain superior cross-task stability and accuracy, highlighting the difficulty of compliance tasks. Furthermore, leveraging our benchmark, we successfully fine-tune a lightweight open-source model. The resulting POLIS series models achieves parity with, or surpasses, strong proprietary baselines on multiple policy subtasks at a significantly reduced cost, providing a cost-effective and compliant path for robust real-world governmental deployment.

View on arXiv PDF

Similar