CVAIJul 10

OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents

arXiv:2607.0906823.2h-index: 9Has Code
Predicted impact top 5% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This benchmark addresses the need for evaluating genuine visual grounding in LVLMs, particularly for map documents, which are a domain where visual information is irreducible to text.

The authors introduce OmniMapBench, a benchmark of 2,096 QA pairs across 1,603 map documents designed to test visual-centric reasoning that cannot be reduced to text. The top-performing LVLM achieves only 75.03% accuracy, highlighting a significant gap in current models.

Recent advancements in LVLMs necessitate robust benchmarks for complex, visually grounded reasoning. A critical limitation is identified in many document understanding benchmarks: visual content is often reducible to text, enabling high performance without genuine visual grounding. To address this limitation, OmniMapBench is introduced to foster visual-centric reasoning for map documents. The benchmark comprises 2,096 manually annotated question-answer pairs across 1,603 map documents from nine categories. It is designed to probe a hierarchy of skills, ranging from perception to multi-step visual reasoning. To quantify benchmark properties, a simple yet effective benchmark-level metric is proposed: the Visual Dependency Index (VDI), defined as the accuracy drop when images are replaced with question-agnostic descriptions. OmniMapBench exhibits higher VDI than established benchmarks, which quantitatively validates its focus on irreducible visual reasoning. Comprehensive evaluations of 25 leading LVLMs are conducted on OmniMapBench. A significant performance gap is observed, with the top-performing model achieving only 75.03\% accuracy. This result underscores the challenges posed by OmniMapBench to current LVLMs. This work aims to catalyze progress in visual-centric reasoning for document understanding of LVLMs. The dataset and code are publicly available at https://github.com/SIGMME/OmniMapBench.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes