REVIEW 10 cited by
GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent advances in large multimodal models (LMMs) have recognized fine-grained grounding as an imperative factor of visual understanding and dialogue. However, the benefits of such representation in LMMs are limited to the natural image domain, and these models perform poorly for remote sensing (RS). The distinct overhead viewpoint, scale variation, and presence of small objects in high-resolution RS imagery present a unique challenge in region-level comprehension. Moreover, the development of the grounding conversation capability of LMMs within RS is hindered by the lack of granular, RS domain-specific grounded data. Addressing these limitations, we propose GeoPixel - the first end-to-end high resolution RS-LMM that supports pixel-level grounding. This capability allows fine-grained visual perception by generating interleaved masks in conversation. GeoPixel supports up to 4K HD resolution in any aspect ratio, ideal for high-precision RS image analysis. To support the grounded conversation generation (GCG) in RS imagery, we curate a visually grounded dataset GeoPixelD through a semi-automated pipeline that utilizes set-of-marks prompting and spatial priors tailored for RS data to methodically control the data generation process. GeoPixel demonstrates superior performance in pixel-level comprehension, surpassing existing LMMs in both single-target and multi-target segmentation tasks. Our methodological ablation studies validate the effectiveness of each component in the overall architecture. Our code and data will be publicly released.
Forward citations
Cited by 10 Pith papers
-
PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation
The work introduces the UAV Reasoning Segmentation task, the DRSeg benchmark dataset, and PixDLM as a baseline dual-path multimodal language model for reasoning-based segmentation in aerial imagery.
-
CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation
CROSS improves remote sensing referring segmentation by combining cascaded SAM distillation with contrastive learning, reporting state-of-the-art cIoU on RefSegRS and RRSIS-D.
-
OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation
OVEarth-Bench, a new open-vocabulary Earth observation benchmark with broad category coverage and diverse queries, shows MLLM-based methods outperform EO-specific ones.
-
GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing
ChronoBench decomposes long-term remote sensing understanding into four cognitive levels, and the GeoChrono model, using per-location temporal trajectories, achieves 78.34% accuracy—over 20 points above prior MLLMs—bu...
-
WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding
A training-free evidence-selection and topology-preserving packaging pipeline improves frozen VLMs on ultra-high-resolution remote sensing VQA without multi-round search.
-
Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning
Aligning images to multi-view caption cores while suppressing orthogonal residual text and disagreement-aware temperature improves robust zero-shot recognition and LVLM transfer.
-
PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis
PathMR applies pixel-level visual reasoning to pathology, jointly generating diagnostic text and cell-type segmentation masks, and introduces the GADVR gastric adenocarcinoma benchmark.
-
GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing
A remote sensing vision-language model that uses task-aware resolution adjustment and attention-based cropping to perform pixel-level segmentation alongside image- and region-level tasks.
-
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
Chain-of-Talkers (CoTalk) has annotators sequentially dictate only the missing visual details, and it reports modest gains in annotation speed and caption density over parallel typed annotation.
-
More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe
An unmodified general-purpose VLM trained with multi-task RL and a SAM3 tool reaches top results on most remote sensing zero-shot benchmarks, with gains the paper attributes to training-data diversity rather than arch...
Discussion (0). Continue with ORCID to comment.