CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia
This work addresses the limitation of current cultural evaluations for LLMs, which often oversimplify culture to factual recall, by providing a more realistic multi-turn simulation for users seeking practical, culturally grounded help in East and Southeast Asia.
This paper introduces CultureConverse, a multilingual, multi-turn simulation and evaluation harness designed to assess large language models' ability to provide culturally grounded assistance across 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains. The resulting dataset, CultureConverse-DS, includes 14,610 benchmark episodes and 274,295 oracle-guided dialogues, with GPT-5 mini achieving the highest assistance quality among 18 models tested.
Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common use case: users seeking practical help over multiple turns in culturally grounded scenarios. We introduce CultureConverse, a scalable, multilingual simulation and evaluation harness for culturally grounded assistant dialogue that covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains. Each simulated and evaluated episode produces a scored interaction where the assistant assists the user and infers cultural constraints from partial information. The resulting CultureConverse-DS dataset contains 14,610 benchmark (evaluation) episodes and 274,295 oracle-guided (gold-mode) dialogues. In our benchmark evaluation of 18 models, GPT-5 mini achieves the highest assistance quality. Human annotation experiments suggest that our evaluation framework is a sufficient proxy for human judgment. Performance gains from fine-tuning on 27,860 high-quality CultureConverse-DS samples improve in-domain assistance and transfer out-of-domain to cultural MCQ and safety classification benchmarks. We release the harness, both splits, and judge prompts to support interactive evaluation of cultural competency.