Pith. sign in

REVIEW 16 cited by

SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.10100 v2 pith:6H2UN2LY submitted 2024-06-14 cs.CV cs.AI

classification cs.CVcs.AI
keywords comprehensionremoterslmmssensingdatasetfit-rsinstructioncomplex
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Remote Sensing Large Multi-Modal Models (RSLMMs) are developing rapidly and showcase significant capabilities in remote sensing imagery (RSI) comprehension. However, due to the limitations of existing datasets, RSLMMs have shortcomings in understanding the rich semantic relations among objects in complex remote sensing scenes. To unlock RSLMMs' complex comprehension ability, we propose a large-scale instruction tuning dataset FIT-RS, containing 1,800,851 instruction samples. FIT-RS covers common interpretation tasks and innovatively introduces several complex comprehension tasks of escalating difficulty, ranging from relation reasoning to image-level scene graph generation. Based on FIT-RS, we build the FIT-RSFG benchmark. Furthermore, we establish a new benchmark to evaluate the fine-grained relation comprehension capabilities of LMMs, named FIT-RSRC. Based on combined instruction data, we propose SkySenseGPT, which achieves outstanding performance on both public datasets and FIT-RSFG, surpassing existing RSLMMs. We hope the FIT-RS dataset can enhance the relation comprehension capability of RSLMMs and provide a large-scale fine-grained data source for the remote sensing community. The dataset will be available at https://github.com/Luo-Z13/SkySenseGPT

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TESSERA v2: Scaling Pixel-wise Earth Foundation Models

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Downstream-driven scaling of pixel-wise Barlow Twins EO models favors large encoders and matched data over projectors, and distillation yields compact Matryoshka students that lead multi-task embedding benchmarks.

  2. NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing

    cs.AI 2026-03 reject novelty 6.5 of 10

    NeSy-Route supplies 10,821 optimally labeled remote-sensing route-planning tasks plus a three-level neuro-symbolic protocol that reveals major perception and planning deficits in current MLLMs.

  3. Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Training on multi-tool visual reasoning trajectories (zoom, grounding, lines) with an attention-focused RL objective improves UHR remote-sensing VQA accuracy over single-tool zoom-in and larger base models.

  4. Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An ordered three-stage post-training route (FBA) improves harbor-scenario performance of RS-MLLMs over direct and collapsed fine-tuning on the authors' HarborEval benchmark.

  5. Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Under one zero-shot protocol, general-purpose MLLMs match or outperform remote-sensing-specific MLLMs on several RS benchmarks, while RS-MLLMs keep advantages in visual grounding, RS-VQA, and ultra-high-resolution und...

  6. Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A new pixel-level aerial streaming referring-segmentation dataset (DroneEyes) and an MLLM (SkyAnchor) with learned token routing and hierarchical memory report SOTA on DroneEyes and large zero-shot gains on SkyFind.

  7. SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SkyVLaM introduces a temporal basis perceiver and adaptive dense selection to improve language-conditioned video segmentation in UAV scenes, and contributes the SkyVid dataset.

  8. GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing

    cs.CV 2026-07 conditional novelty 6.0 of 10

    ChronoBench decomposes long-term remote sensing understanding into four cognitive levels, and the GeoChrono model, using per-location temporal trajectories, achieves 78.34% accuracy—over 20 points above prior MLLMs—bu...

  9. OSMDA: OpenStreetMap-based Domain Adaptation for Remote Sensing VLMs

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A remote-sensing VLM can be adapted by having it read rendered OpenStreetMap maps paired with satellite images, then fine-tuning it on satellite images alone.

  10. OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A 4B model fine-tuned on tool-augmented geospatial reasoning traces outperforms larger general-purpose models on executable GIS/spectral tool-use benchmarks and matches frontier models on trajectory fidelity.

  11. VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    VectorLLM, a multimodal LLM that regresses building contour vertices token by token, reports gains of 5.6 to 13.6 AP over prior polygon extraction methods on WHU, WHU-Mix, and CrowdAI.

  12. GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GeoProg3D combines a georeferenced hierarchical 3D language field, geographic vision APIs, and LLM-generated programs to answer natural-language queries about city-scale 3D scenes, and includes a new 952-query benchma...

  13. UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A fine-tuned small multimodal LLM outperforms much larger general models on urban tasks in a new benchmark, with caveats about benchmark overlap with training data.

  14. More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

    cs.CV 2026-07 conditional novelty 5.0 of 10

    An unmodified general-purpose VLM trained with multi-task RL and a SAM3 tool reaches top results on most remote sensing zero-shot benchmarks, with gains the paper attributes to training-data diversity rather than arch...

  15. Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A remote sensing LVLM that augments visual features with retrieved captions and routes them through level-specific experts improves performance on several RS vision-language benchmarks.

  16. Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives

    cs.CV 2025-05 conditional novelty 3.0 of 10

    A structured review of remote sensing vision-language models, organizing contrastive, instruction-tuned, and generative approaches alongside their datasets and benchmarks.

Pith tools