REVIEW 9 cited by
RSGPT: A Remote Sensing Vision Language Model and Benchmark
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The emergence of large-scale large language models, with GPT-4 as a prominent example, has significantly propelled the rapid advancement of artificial general intelligence and sparked the revolution of Artificial Intelligence 2.0. In the realm of remote sensing (RS), there is a growing interest in developing large vision language models (VLMs) specifically tailored for data analysis in this domain. However, current research predominantly revolves around visual recognition tasks, lacking comprehensive, large-scale image-text datasets that are aligned and suitable for training large VLMs, which poses significant challenges to effectively training such models for RS applications. In computer vision, recent research has demonstrated that fine-tuning large vision language models on small-scale, high-quality datasets can yield impressive performance in visual and language understanding. These results are comparable to state-of-the-art VLMs trained from scratch on massive amounts of data, such as GPT-4. Inspired by this captivating idea, in this work, we build a high-quality Remote Sensing Image Captioning dataset (RSICap) that facilitates the development of large VLMs in the RS field. Unlike previous RS datasets that either employ model-generated captions or short descriptions, RSICap comprises 2,585 human-annotated captions with rich and high-quality information. This dataset offers detailed descriptions for each image, encompassing scene descriptions (e.g., residential area, airport, or farmland) as well as object information (e.g., color, shape, quantity, absolute position, etc). To facilitate the evaluation of VLMs in the field of RS, we also provide a benchmark evaluation dataset called RSIEval. This dataset consists of human-annotated captions and visual question-answer pairs, allowing for a comprehensive assessment of VLMs in the context of RS.
Forward citations
Cited by 9 Pith papers
-
WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding
A training-free evidence-selection and topology-preserving packaging pipeline improves frozen VLMs on ultra-high-resolution remote sensing VQA without multi-round search.
-
Few-Shot Vision-Language Reasoning for Satellite Imagery via Verifiable Rewards
A few-shot RLVR method using only rule-based rewards lifts a 2B vision-language model's remote sensing accuracy by double digits, with 128 examples rivaling thousands.
-
Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation
A two-stage method (MpGI) produces a 210K-image, 1.26M-caption remote sensing dataset and state-of-the-art CLIP and CoCa models.
-
GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing
A remote sensing vision-language model that uses task-aware resolution adjustment and attention-based cropping to perform pixel-level segmentation alongside image- and region-level tasks.
-
A Satellite-Ground Synergistic Large Vision-Language Model System for Earth Observation
SpaceVerse jointly decides where to run vision-language inference in LEO satellite networks and compresses task-irrelevant image regions before downlink, improving accuracy and cutting latency versus baselines.
-
GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields
GeoProg3D combines a georeferenced hierarchical 3D language field, geographic vision APIs, and LLM-generated programs to answer natural-language queries about city-scale 3D scenes, and includes a new 952-query benchma...
-
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
A fine-tuned small multimodal LLM outperforms much larger general models on urban tasks in a new benchmark, with caveats about benchmark overlap with training data.
-
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
Chain-of-Talkers (CoTalk) has annotators sequentially dictate only the missing visual details, and it reports modest gains in annotation speed and caption density over parallel typed annotation.
-
SARChat-Bench-2M: A Multi-Task Vision-Language Benchmark for SAR Image Interpretation
SARChat-2M is a 2M-sample, six-task instruction-tuning dataset and benchmark for vision-language models on synthetic aperture radar imagery.
Discussion (0). Continue with ORCID to comment.