REVIEW 5 cited by
MMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
MMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation Learning
read the original abstract
The volume of unlabelled Earth observation (EO) data is huge, but many important applications lack labelled training data. However, EO data offers the unique opportunity to pair data from different modalities and sensors automatically based on geographic location and time, at virtually no human labor cost. We seize this opportunity to create MMEarth, a diverse multi-modal pretraining dataset at global scale. Using this new corpus of 1.2 million locations, we propose a Multi-Pretext Masked Autoencoder (MP-MAE) approach to learn general-purpose representations for optical satellite images. Our approach builds on the ConvNeXt V2 architecture, a fully convolutional masked autoencoder (MAE). Drawing upon a suite of multi-modal pretext tasks, we demonstrate that our MP-MAE approach outperforms both MAEs pretrained on ImageNet and MAEs pretrained on domain-specific satellite images. This is shown on several downstream tasks including image classification and semantic segmentation. We find that pretraining with multi-modal pretext tasks notably improves the linear probing performance compared to pretraining on optical satellite images only. This also leads to better label efficiency and parameter efficiency which are crucial aspects in global scale applications.
Forward citations
Cited by 5 Pith papers
-
SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm
SARLO-80 is a new public dataset of 119566 complex SAR-optical-text triplets standardized to 80cm slant-range resolution from 257 locations across 72 countries.
-
SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm
SARLO-80 provides 119,566 complex SAR-optical-text triplets at 80 cm slant-range resolution with fixed splits and preprocessing code.
-
Cryo-Bench: Benchmarking Foundation Models for Cryosphere Applications
A new benchmark shows a simple U-Net outperforms frozen geospatial foundation models on cryosphere segmentation, but fine-tuning with learning-rate tuning narrows or reverses the gap.
-
On What Depends the Robustness of Multi-source Models to Missing Data in Earth Observation?
Multi-source EO models' robustness to missing data depends on task nature, source complementarity, and design, sometimes improving when certain sources are removed.
-
Scalable and Trustworthy Earth Observation Foundation Models
Remote-sensing foundation models need domain-specific design and evaluation around measurement physics and decision constraints; benchmark accuracy alone is insufficient for trustworthy EO deployment.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.