CLAIIRJun 17

TW-LegalBench: Measuring Taiwanese Legal Understanding

arXiv:2606.1869918.3
Predicted impact top 49% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and practitioners in legal AI, this benchmark fills a gap in civil-law evaluation for Traditional Chinese, revealing that while LLMs approach human-level performance on qualification exams, reliable legal text generation remains challenging.

TW-LegalBench evaluates LLMs on Taiwanese legal reasoning using over 30,000 instances across multiple tasks. Top models exceed the lawyer qualification exam passing rate (11%) but fall short of the judge/prosecutor rate (1-2%), and struggle with legal article citation.

Large language models (LLMs) have shown impressive capabilities across diverse tasks, yet their performance on jurisdiction-specific legal reasoning remains underexplored. We present TW-LegalBench that utilizes Taiwanese legal system's rich official corpus open to the public to fill the gap in evaluating LLMs on Taiwanese law, among common-law benchmarks that focus on English sources and civil-law benchmarks focusing on sources of Simplified Chinese. TW-LegalBench comprises three task types: (1) over 16,000 multiple-choice questions (MCQs) across five years of official examinations in 18 professional domains; (2) 117 open-ended essay questions (OEQs) from examinations for legal professionals with official scoring rubrics; and (3) more than 14,000 legal judgment prediction (LJP) instances covering hundreds of crime categories. We evaluate 13 LLMs using accuracy for MCQs, a decomposed LLM-as-Judge framework based on the scoring rubric points for OEQs, and metrics for sentencing accuracy and statute citation for LJP. Our results reveal that top-performing models exceed the passing threshold for qualified lawyers (passing rate: 11%) but fall short of that for judges and prosecutors (passing rate: 1~2%). For LJP, while models demonstrate reasonable verdict type accuracy and sentence prediction capability, they struggle to cite exact legal articles. These findings highlight that reliable legal text generation remains challenging for LLMs, even though their performance on qualification examinations approaches human level.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes