CLJun 19

CulMind: Benchmarking Multimodal Understanding and Reasoning in Chinese Cultural Heritage

arXiv:2606.2161822.2Has Code
Predicted impact top 30% in CL · last 90 daysOriginality Incremental advance
AI Analysis

This work provides a more nuanced evaluation framework for multimodal reasoning in a specific cultural domain, addressing the underexplored aspect of reasoning process quality.

The authors introduce CulMind and CulMind-R, benchmarks for multimodal understanding and reasoning in Chinese Cultural Heritage, along with ReaScore, a task-adaptive metric for evaluating reasoning quality. Experiments on 14 MLLMs show a significant gap between answer accuracy and reasoning quality, especially on challenging tasks.

Evaluating Multimodal Large Language Models (MLLMs) in Chinese Cultural Heritage (CCH) requires fine-grained reasoning over visual, textual, stylistic, and historical clues. However, existing CCH benchmarks mainly emphasize final-answer accuracy, while the accuracy and completeness of reasoning processes remain underexplored. To address this gap, we introduce CulMind and CulMind-R: a high-quality benchmark for multimodal CCH covering 50 tasks from collections of more than 100 museums, and a 24-task reasoning subset that adaptively defines task-specific dimensions for reasoning process evaluation. To evaluate reasoning quality, we propose ReaScore, a task-adaptive metric that evaluates reasoning by automatically weighting task-relevant dimensions. Experiments on 14 leading MLLMs reveal a substantial gap between answers and reasoning, especially on challenging tasks. Further analysis shows that task-adaptive dimension selection and weighting better align evaluation results with expert judgments. Overall, our benchmark and metric support a more expert-aligned assessment of CCH understanding and offer a transferable reference for broader evaluations of cultural heritage. We publicly release the data, code, and evaluation scripts at https://github.com/ZevTsao/CulMind to facilitate reproducible research.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes