REVIEW 4 cited by
Vision Foundation Models in Remote Sensing: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Artificial Intelligence (AI) technologies have profoundly transformed the field of remote sensing, revolutionizing data collection, processing, and analysis. Traditionally reliant on manual interpretation and task-specific models, remote sensing research has been significantly enhanced by the advent of foundation models-large-scale, pre-trained AI models capable of performing a wide array of tasks with unprecedented accuracy and efficiency. This paper provides a comprehensive survey of foundation models in the remote sensing domain. We categorize these models based on their architectures, pre-training datasets, and methodologies. Through detailed performance comparisons, we highlight emerging trends and the significant advancements achieved by those foundation models. Additionally, we discuss technical challenges, practical implications, and future research directions, addressing the need for high-quality data, computational resources, and improved model generalization. Our research also finds that pre-training methods, particularly self-supervised learning techniques like contrastive learning and masked autoencoders, remarkably enhance the performance and robustness of foundation models. This survey aims to serve as a resource for researchers and practitioners by providing a panorama of advances and promising pathways for continued development and application of foundation models in remote sensing.
Forward citations
Cited by 4 Pith papers
-
MAPEX: Modality-Aware Pruning of Experts for Remote Sensing Foundation Models
MAPEX shows that a modality-conditioned mixture-of-experts vision transformer, pre-trained on six remote sensing modalities and then pruned to keep only the experts for a target modality, can outperform or match large...
-
Time2Agri: Temporal Pretext Tasks for Agricultural Monitoring
Temporal pretext tasks (time-difference, frequency, future-frame prediction) improve SSL representations for crop mapping and yield estimation, with future-frame prediction best on SICKLE and FTW India.
-
Leveraging Satellite Image Time Series for Accurate Extreme Event Detection
SITS-Extreme detects extreme events by learning patch representations from a satellite image time series via autoencoding with contrastive and consistency losses, then thresholding the mean cosine distance between pre...
-
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
RDB improves remote sensing image-text retrieval mean recall by 1.15 to 2 percent over fully fine-tuned GeoRSCLIP using an asymmetric adapter and a dual-task consistency loss.
Discussion (0). Continue with ORCID to comment.