REVIEW 4 cited by
USat: A Unified Self-Supervised Encoder for Multi-Sensor Satellite Imagery
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large, self-supervised vision models have led to substantial advancements for automatically interpreting natural images. Recent works have begun tailoring these methods to remote sensing data which has rich structure with multi-sensor, multi-spectral, and temporal information providing massive amounts of self-labeled data that can be used for self-supervised pre-training. In this work, we develop a new encoder architecture called USat that can input multi-spectral data from multiple sensors for self-supervised pre-training. USat is a vision transformer with modified patch projection layers and positional encodings to model spectral bands with varying spatial scales from multiple sensors. We integrate USat into a Masked Autoencoder (MAE) self-supervised pre-training procedure and find that a pre-trained USat outperforms state-of-the-art self-supervised MAE models trained on remote sensing data on multiple remote sensing benchmark datasets (up to 8%) and leads to improvements in low data regimes (up to 7%). Code and pre-trained weights are available at https://github.com/stanfordmlgroup/USat .
Forward citations
Cited by 4 Pith papers
-
AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities
A single JEPA-based model with scale-adaptive encoders is pre-trained on five heterogeneous Earth observation datasets and reaches state-of-the-art results across nine downstream tasks.
-
SatMamba: Development of Foundation Models for Remote Sensing Imagery Using State Space Models
SatMamba shows that a Mamba-based masked autoencoder matches ViT-based MAE on remote sensing segmentation and damage assessment, with efficiency linear in sequence length only at larger inputs.
-
Deploying Geospatial Foundation Models in the Real World: Lessons from WorldCereal
A structured protocol for deploying geospatial foundation models is introduced and validated in WorldCereal, where fine-tuned Presto outperforms a fully-supervised CatBoost baseline in crop mapping.
-
TiMo: Spatiotemporal Foundation Model for Satellite Image Time Series
TiMo, a hierarchical transformer pretrained on one million Sentinel-2 images with a space-time gyroscope attention, reports state-of-the-art accuracy on deforestation, land cover, crop type, and flood mapping tasks.
Discussion (0). Continue with ORCID to comment.