REVIEW 3 cited by
Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce a method to train vision-language models for remote-sensing images without using any textual annotations. Our key insight is to use co-located internet imagery taken on the ground as an intermediary for connecting remote-sensing images and language. Specifically, we train an image encoder for remote sensing images to align with the image encoder of CLIP using a large amount of paired internet and satellite images. Our unsupervised approach enables the training of a first-of-its-kind large-scale vision language model (VLM) for remote sensing images at two different resolutions. We show that these VLMs enable zero-shot, open-vocabulary image classification, retrieval, segmentation and visual question answering for satellite images. On each of these tasks, our VLM trained without textual annotations outperforms existing VLMs trained with supervision, with gains of up to 20% for classification and 80% for segmentation.
Forward citations
Cited by 3 Pith papers
-
DiSciPLE: Learning Interpretable Programs for Scientific Visual Discovery
DiSciPLE uses LLM-guided evolution to discover interpretable Python programs that predict geospatial quantities, outperforming black-box deep nets on population density and on out-of-distribution generalization.
-
Advancing ALS Applications with Large-Scale Pre-training: Dataset Development and Downstream Assessment
A new large-scale ALS point cloud pre-training dataset, sampled by land cover and slope diversity, improves downstream task performance when used to pre-train BEV-MAE.
-
REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation
A vision-language model for Earth observation that adds numeric biomass regression to generative question answering, with R^2 up to 0.36 and patch counting still failing.
Discussion (0). Continue with ORCID to comment.