CLAIOct 17, 2025

KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models

arXiv:2510.15558v1h-index: 13
Originality Synthesis-oriented
AI Analysis

This addresses the problem of linguistic and cultural bias in LLM evaluations for researchers and developers focusing on Korean and underrepresented languages, though it is incremental as it extends existing benchmarking approaches to a new language.

The paper tackles the lack of benchmarks for evaluating instruction-following abilities in large language models for Korean, introducing KITE to assess open-ended tasks and revealing performance disparities across models through automated and human evaluations.

The instruction-following capabilities of large language models (LLMs) are pivotal for numerous applications, from conversational agents to complex reasoning systems. However, current evaluations predominantly focus on English models, neglecting the linguistic and cultural nuances of other languages. Specifically, Korean, with its distinct syntax, rich morphological features, honorific system, and dual numbering systems, lacks a dedicated benchmark for assessing open-ended instruction-following capabilities. To address this gap, we introduce the Korean Instruction-following Task Evaluation (KITE), a comprehensive benchmark designed to evaluate both general and Korean-specific instructions. Unlike existing Korean benchmarks that focus mainly on factual knowledge or multiple-choice testing, KITE directly targets diverse, open-ended instruction-following tasks. Our evaluation pipeline combines automated metrics with human assessments, revealing performance disparities across models and providing deeper insights into their strengths and weaknesses. By publicly releasing the KITE dataset and code, we aim to foster further research on culturally and linguistically inclusive LLM development and inspire similar endeavors for other underrepresented languages.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes