CVJul 22

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

arXiv:2607.2028413.8
Predicted impact top 18% in CV · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers and practitioners in remote sensing, this work clarifies the capability boundaries of MLLMs, showing that general-purpose models are competitive and highlighting remaining challenges.

This survey evaluates multimodal large language models (MLLMs) for remote sensing image understanding, finding that general-purpose computer vision MLLMs can match or outperform domain-specific remote sensing MLLMs on several tasks without fine-tuning, while both types face limitations in spatial reasoning and fine-grained understanding.

The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery. However, a systematic understanding of the capability boundaries, cross-task generalization, and task-specific limitations of existing remote sensing MLLMs (RS-MLLMs) is still lacking. This paper presents a systematic survey and diagnostic evaluation of MLLMs for RSISU. We review the technical evolution of RS-MLLMs, focusing on model design, multimodal learning, training data, and downstream capabilities. We further compare RS-MLLMs with general-purpose computer vision MLLMs (CV-MLLMs) across diverse RSISU tasks and benchmarks. RS-MLLMs remain competitive in domain-specific settings, particularly remote sensing visual grounding and high-resolution visual question answering. More notably, general-purpose CV-MLLMs can match or even outperform these specialized models on several RSISU tasks without remote sensing-specific fine-tuning. These findings demonstrate the strong transferability of general-purpose CV-MLLMs and show that current RS-MLLMs do not consistently outperform them across diverse RSISU tasks. Current MLLMs also face limitations in spatial and relational reasoning, fine-grained visual understanding, instruction diversity, and generalization across heterogeneous task formats. Based on these findings, we outline future directions toward reliable evaluation, multimodal and high-resolution reasoning, efficient deployment, and tool-augmented remote sensing agents. This survey provides a systematic reference for developing robust, generalizable, and practical MLLMs for RSISU.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes