Pith. sign in

REVIEW 5 cited by

AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.14123 v3 pith:L2G3AG4H submitted 2024-12-18 cs.CV

classification cs.CV
keywords modelanysatdatasetsclassificationdataearthgeoplexmodalities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Geospatial models must adapt to the diversity of Earth observation data in terms of resolutions, scales, and modalities. However, existing approaches expect fixed input configurations, which limits their practical applicability. We propose AnySat, a multimodal model based on joint embedding predictive architecture (JEPA) and scale-adaptive spatial encoders, allowing us to train a single model on highly heterogeneous data in a self-supervised manner. To demonstrate the advantages of this unified approach, we compile GeoPlex, a collection of 5 multimodal datasets with varying characteristics and $11$ distinct sensors. We then train a single powerful model on these diverse datasets simultaneously. Once fine-tuned or probed, we reach state-of-the-art results on the test sets of GeoPlex and for 6 external datasets across various environment monitoring tasks: land cover mapping, tree species identification, crop type classification, change detection, climate type classification, and segmentation of flood, burn scar, and deforestation. The code and models are available at https://github.com/gastruc/AnySat.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. TESSERA v2: Scaling Pixel-wise Earth Foundation Models

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Downstream-driven scaling of pixel-wise Barlow Twins EO models favors large encoders and matched data over projectors, and distillation yields compact Matryoshka students that lead multi-task embedding benchmarks.

  2. Emerging Flexible Designs for Geospatial Multimodal Foundation Models

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    Standardized pretraining and evaluation of geospatial multimodal foundation models on GEOBench reveals design trade-offs in flexibility, modality alignment, and task performance.

  3. MOMO: Mars Orbital Model Foundation Model for Mars Orbital Applications

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    MOMO merges sensor-specific models from three Mars orbital instruments at matched validation loss stages to form a foundation model that outperforms ImageNet, Earth observation, sensor-specific, and supervised baselin...

  4. Large scale cross-regional remote sensing flood monitoring framework for operative mapping and impact analysis

    cs.CV 2026-07 conditional novelty 4.0 of 10

    Under limited Russian flood labels, multimodal U-Net++ (F1 0.84) outperforms fine-tuned AnySat for water mapping, and the masks feed an EMERCOM-style damage pipeline that matches Tulun 2019 official area and exposure ...

  5. Scalable and Trustworthy Earth Observation Foundation Models

    cs.LG 2026-07 conditional novelty 3.0 of 10

    Remote-sensing foundation models need domain-specific design and evaluation around measurement physics and decision constraints; benchmark accuracy alone is insufficient for trustworthy EO deployment.

Pith tools