Pith. sign in

REVIEW 14 cited by

LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.02544 v4 pith:NDSONT75 submitted 2024-02-04 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords languagelargelhrs-botunderstandingdatasetdiverseimageimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The revolutionary capabilities of large language models (LLMs) have paved the way for multimodal large language models (MLLMs) and fostered diverse applications across various specialized domains. In the remote sensing (RS) field, however, the diverse geographical landscapes and varied objects in RS imagery are not adequately considered in recent MLLM endeavors. To bridge this gap, we construct a large-scale RS image-text dataset, LHRS-Align, and an informative RS-specific instruction dataset, LHRS-Instruct, leveraging the extensive volunteered geographic information (VGI) and globally available RS images. Building on this foundation, we introduce LHRS-Bot, an MLLM tailored for RS image understanding through a novel multi-level vision-language alignment strategy and a curriculum learning method. Additionally, we introduce LHRS-Bench, a benchmark for thoroughly evaluating MLLMs' abilities in RS image understanding. Comprehensive experiments demonstrate that LHRS-Bot exhibits a profound understanding of RS images and the ability to perform nuanced reasoning within the RS domain.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    VectorLLM, a multimodal LLM that regresses building contour vertices token by token, reports gains of 5.6 to 13.6 AP over prior polygon extraction methods on WHU, WHU-Mix, and CrowdAI.

  2. UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A fine-tuned small multimodal LLM outperforms much larger general models on urban tasks in a new benchmark, with caveats about benchmark overlap with training data.

  3. ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    ZeroBench is a hand-built 100-question visual reasoning benchmark, adversarially filtered so every evaluated frontier LMM scored 0% at release.

  4. GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing

    cs.CV 2025-01 conditional novelty 6.0 of 10

    GeoPixel brings pixel-level grounding to remote sensing large multimodal models for the first time, with a new dataset and benchmark built from iSAID.

  5. GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

    cs.CV 2025-01 conditional novelty 6.0 of 10

    GeoPix extends multimodal language models for remote sensing to pixel-level referring segmentation through a mask predictor, a class-wise learnable memory, and a new 65,463-image instruction dataset.

  6. Cross-View Image Set Geo-Localization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Using several unordered ground-view photos as a query set improves cross-view geo-localization accuracy, and the proposed FlexGeo model achieves state-of-the-art results on a new six-city benchmark.

  7. EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues

    cs.CV 2024-12 conditional novelty 6.0 of 10

    EarthDial is a 4B-parameter remote sensing chatbot trained on 11.11M instruction pairs to handle multi-resolution, multi-spectral, and multi-temporal satellite imagery, and it reports gains over prior VLMs on dozens o...

  8. RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts

    cs.CV 2024-12 conditional novelty 6.0 of 10

    RSUniVLM is a 1-billion-parameter remote sensing vision-language model that unifies image-, region-, and pixel-level tasks plus multi-image change analysis, achieving state-of-the-art visual grounding on VRSBench and ...

  9. UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A single vision-language model is fine-tuned to handle three remote sensing input types and reports state-of-the-art results on VQA, change captioning, and video classification benchmarks.

  10. GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding

    cs.CV 2024-11 conditional novelty 5.0 of 10

    GeoGround unifies horizontal box, oriented box, and mask visual grounding in a single remote sensing vision-language model using a text-based mask representation.

  11. LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A remote sensing chatbot with recaptioned image-text data and a mixture-of-experts visual bridge reports gains over prior general and remote sensing models, but no code or data are released yet.

  12. A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A survey of 80+ Deep Research systems that proposes a four-layer taxonomy (foundation models, tool use, planning, synthesis) and compares commercial and open-source implementations.

  13. REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation

    cs.CV 2024-12 reject novelty 4.0 of 10

    A vision-language model for Earth observation that adds numeric biomass regression to generative question answering, with R^2 up to 0.36 and patch counting still failing.

  14. Visual Large Language Models for Generalized and Specialized Applications

    cs.CV 2025-01 conditional novelty 3.0 of 10

    This paper reviews and taxonomizes VLLM applications into vision-to-text, vision-to-action, and text-to-vision, adding ethics and future-work discussion.

Pith tools