Pith. sign in

REVIEW 5 cited by

MMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.02771 v2 pith:IXM7KTJK submitted 2024-05-04 cs.CV cs.AIcs.LG

MMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation Learning

classification cs.CV cs.AIcs.LG
keywords datamulti-modaltasksapproachimagespretextpretrainingsatellite
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The volume of unlabelled Earth observation (EO) data is huge, but many important applications lack labelled training data. However, EO data offers the unique opportunity to pair data from different modalities and sensors automatically based on geographic location and time, at virtually no human labor cost. We seize this opportunity to create MMEarth, a diverse multi-modal pretraining dataset at global scale. Using this new corpus of 1.2 million locations, we propose a Multi-Pretext Masked Autoencoder (MP-MAE) approach to learn general-purpose representations for optical satellite images. Our approach builds on the ConvNeXt V2 architecture, a fully convolutional masked autoencoder (MAE). Drawing upon a suite of multi-modal pretext tasks, we demonstrate that our MP-MAE approach outperforms both MAEs pretrained on ImageNet and MAEs pretrained on domain-specific satellite images. This is shown on several downstream tasks including image classification and semantic segmentation. We find that pretraining with multi-modal pretext tasks notably improves the linear probing performance compared to pretraining on optical satellite images only. This also leads to better label efficiency and parameter efficiency which are crucial aspects in global scale applications.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm

    cs.CV 2026-06 accept novelty 8.0

    SARLO-80 is a new public dataset of 119566 complex SAR-optical-text triplets standardized to 80cm slant-range resolution from 257 locations across 72 countries.

  2. SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm

    cs.CV 2026-06 conditional novelty 7.0

    SARLO-80 provides 119,566 complex SAR-optical-text triplets at 80 cm slant-range resolution with fixed splits and preprocessing code.

  3. Cryo-Bench: Benchmarking Foundation Models for Cryosphere Applications

    cs.CV 2026-03 conditional novelty 6.0

    A new benchmark shows a simple U-Net outperforms frozen geospatial foundation models on cryosphere segmentation, but fine-tuning with learning-rate tuning narrows or reverses the gap.

  4. On What Depends the Robustness of Multi-source Models to Missing Data in Earth Observation?

    cs.LG 2025-03 unverdicted novelty 4.0

    Multi-source EO models' robustness to missing data depends on task nature, source complementarity, and design, sometimes improving when certain sources are removed.

  5. Scalable and Trustworthy Earth Observation Foundation Models

    cs.LG 2026-07 conditional novelty 3.0

    Remote-sensing foundation models need domain-specific design and evaluation around measurement physics and decision constraints; benchmark accuracy alone is insufficient for trustworthy EO deployment.