CVJun 30

Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding

arXiv:2606.3116915.2
Predicted impact top 17% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in AI-assisted paleography, this benchmark reveals that current MLLMs cannot handle structured inscription-level understanding, highlighting a key limitation.

This work introduces S-OBI, a benchmark for evaluating multimodal large language models (MLLMs) on sentence-level oracle bone inscription understanding. Experiments show that current MLLMs perform poorly, with visual perception errors in unmasked regions propagating to cause erroneous predictions for masked characters, indicating strong dependence on character-level recognition.

Existing AI-assisted oracle bone inscription (OBI) visual recognition and understanding studies mainly focus on character-level, ignoring the long-form textual coherence and contextual dependencies embedded in complete divination charges. Recently, the powerful visual perception capabilities of multimodal large language models (MLLMs) have opened new possibilities for OBI information processing. In this work, we introduce S-OBI, a novel benchmark for evaluating MLLMs in Sentence-level OBI understanding. Instead of using noisy and incomplete rubbings as the visual input, S-OBI synthesizes clear and standardized sentence-level OBI instances through glyph substitution and composition. According to 95 original rubbings with translations that have been identified, corrected, and verified by experts, we replace characters in the original rubbings with corresponding clean glyph samples sourced from existing OBI datasets while preserving the overall inscriptional structure and semantic organization. This mitigates the influence of low-level distortions and enables a more focused evaluation of sentence-level OBI understanding. Based on this, we design semantic matching, semantic slot extraction, and contextual reasoning tasks and obtain 695 question-answer pairs. Experiments reveal the inferiority of contemporary MLLMs on sentence-level OBI understanding. In particular, visual perception errors in unmasked regions propagate through the reasoning chain, leading to erroneous predictions for masked characters, which indicates that sentence-level OBI understanding in current models remains strongly dependent on character-level recognition. Overall, S-OBI provides a diagnostic benchmark for evaluating whether MLLMs can move beyond isolated character recognition toward structured inscription-level understanding.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes