Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?
Merged summary
TL;DR - This survey and diagnostic evaluation compares remote-sensing-specific MLLMs with general-purpose vision-language models. General-purpose models often match or outperform specialized systems, highlighting strong transferability and gaps in current domain-specific approaches.
- RS-MLLMs remain competitive in visual grounding and high-resolution visual question answering.
- General-purpose CV-MLLMs perform strongly without remote-sensing-specific fine-tuning.
- Key limitations include spatial and relational reasoning, fine-grained perception, instruction diversity, and cross-task generalization.
- Future priorities include reliable evaluation, high-resolution reasoning, efficient deployment, and tool-augmented agents.
Sources (1)
Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?
TL;DR - This survey and diagnostic evaluation compares remote-sensing-specific MLLMs with general-purpose vision-language models. General-purpose models often match or outperform specialized systems, highlighting strong transferability and gaps in current domain-specific approaches.
- RS-MLLMs remain competitive in visual grounding and high-resolution visual question answering.
- General-purpose CV-MLLMs perform strongly without remote-sensing-specific fine-tuning.
- Key limitations include spatial and relational reasoning, fine-grained perception, instruction diversity, and cross-task generalization.
- Future priorities include reliable evaluation, high-resolution reasoning, efficient deployment, and tool-augmented agents.