Pith. sign in

REVIEW 1 cited by

Pushing the Limits of Vision-Language Models in Remote Sensing without Human Annotations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.07048 v1 pith:M34OYQLS submitted 2024-09-11 cs.CV

classification cs.CV
keywords vision-languagedatasetsmodelsfoundationmodelclassificationdomainemploying
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The prominence of generalized foundation models in vision-language integration has witnessed a surge, given their multifarious applications. Within the natural domain, the procurement of vision-language datasets to construct these foundation models is facilitated by their abundant availability and the ease of web crawling. Conversely, in the remote sensing domain, although vision-language datasets exist, their volume is suboptimal for constructing robust foundation models. This study introduces an approach to curate vision-language datasets by employing an image decoding machine learning model, negating the need for human-annotated labels. Utilizing this methodology, we amassed approximately 9.6 million vision-language paired datasets in VHR imagery. The resultant model outperformed counterparts that did not leverage publicly available vision-language datasets, particularly in downstream tasks such as zero-shot classification, semantic localization, and image-text retrieval. Moreover, in tasks exclusively employing vision encoders, such as linear probing and k-NN classification, our model demonstrated superior efficacy compared to those relying on domain-specific vision-language datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval

    cs.CV 2025-05 conditional novelty 4.0 of 10

    RDB improves remote sensing image-text retrieval mean recall by 1.15 to 2 percent over fully fine-tuned GeoRSCLIP using an asymmetric adapter and a dual-task consistency loss.

Pith tools