Pith. sign in

REVIEW 25 cited by

PANGAEA: A Global and Inclusive Benchmark for Geospatial Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.04204 v2 pith:PIMZYXLC submitted 2024-12-05 cs.CV

PANGAEA: A Global and Inclusive Benchmark for Geospatial Foundation Models

classification cs.CV
keywords gfmsmodelsevaluationbenchmarkdatasetsgeospatialpangaeatasks
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Geospatial Foundation Models (GFMs) have emerged as powerful tools for extracting representations from Earth observation data, but their evaluation remains inconsistent and narrow. Existing works often evaluate on suboptimal downstream datasets and tasks, that are often too easy or too narrow, limiting the usefulness of the evaluations to assess the real-world applicability of GFMs. Additionally, there is a distinct lack of diversity in current evaluation protocols, which fail to account for the multiplicity of image resolutions, sensor types, and temporalities, which further complicates the assessment of GFM performance. In particular, most existing benchmarks are geographically biased towards North America and Europe, questioning the global applicability of GFMs. To overcome these challenges, we introduce PANGAEA, a standardized evaluation protocol that covers a diverse set of datasets, tasks, resolutions, sensor modalities, and temporalities. It establishes a robust and widely applicable benchmark for GFMs. We evaluate the most popular GFMs openly available on this benchmark and analyze their performance across several domains. In particular, we compare these models to supervised baselines (e.g. UNet and vanilla ViT), and assess their effectiveness when faced with limited labeled data. Our findings highlight the limitations of GFMs, under different scenarios, showing that they do not consistently outperform supervised models. PANGAEA is designed to be highly extensible, allowing for the seamless inclusion of new datasets, models, and tasks in future research. By releasing the evaluation code and benchmark, we aim to enable other researchers to replicate our experiments and build upon our work, fostering a more principled evaluation protocol for large pre-trained geospatial models. The code is available at https://github.com/VMarsocci/pangaea-bench.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Moonstone: A Multimodal Foundation Model and Benchmark for Lunar Remote Sensing

    cs.CV 2026-07 accept novelty 7.0

    MG-MAE pretrained on a new 28-channel lunar map outperforms ImageNet, vanilla MAE, and EO foundation models across six classification, regression, and segmentation tasks.

  2. Uncertainty-aware tree height change regression

    cs.CV 2026-07 unverdicted novelty 7.0

    Introduces the CHC dataset of 3 m continuous canopy height differences with uncertainties and the uncertainty-aware change regression task for fine-tuning GFMs on PlanetScope time series.

  3. UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observation

    cs.CV 2026-06 unverdicted novelty 7.0

    UniverSat is a ViT-style model with a universal patch encoder enabling self-supervised training on heterogeneous multimodal Earth observation data from varying resolutions and sensors.

  4. GeoNatureAgent Benchmark: Benchmarking LLM Agents for Environmental Geospatial Analysis Across Frontier and Open-Weight Foundation Models

    cs.AI 2026-06 accept novelty 7.0

    GeoNatureAgent Benchmark tests seven LLMs on 93 tasks via a production geospatial API, with Claude Sonnet 4 at 60.8% and DeepSeek V3.2 offering near performance at 11x lower cost while all models fail on close-value c...

  5. SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals

    cs.CV 2026-05 unverdicted novelty 7.0

    SDGBiasBench reveals intrinsic SDG biases in VLMs driven by priors rather than evidence, and CADE mitigates them with up to 25% accuracy gains and 12-point MAE reductions.

  6. Does Your Wildfire Prediction Model Actually Work, or Just Score Well?

    cs.LG 2026-05 unverdicted novelty 7.0

    WILDFIRE-FM is the first wildfire-specific Earth foundation model, paired with a fixed-contract evaluation framework that demonstrates wildfire model transfer conclusions depend strongly on evaluation design and task ...

  7. NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing

    cs.AI 2026-03 accept novelty 6.5

    NeSy-Route supplies 10,821 optimally labeled remote-sensing route-planning tasks plus a three-level neuro-symbolic protocol that reveals major perception and planning deficits in current MLLMs.

  8. Does Your Wildfire Prediction Model Actually Work, or Just Score Well?

    cs.LG 2026-05 unverdicted novelty 6.0

    Introduces WILDFIRE-FM and a fixed-contract evaluation framework demonstrating that wildfire model transfer conclusions depend strongly on evaluation design and task formulation.

  9. No One Knows the State of the Art in Geospatial Foundation Models

    cs.CV 2026-05 accept novelty 6.0

    An audit of 152 papers reveals that geospatial foundation models lack standardized evaluations, training controls, and weight releases, so no one knows the state of the art.

  10. LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation

    cs.CV 2026-05 conditional novelty 6.0

    LithoBench is a new multi-level benchmark showing that existing large multimodal models have substantial limitations in geological semantic understanding for remote sensing lithology interpretation.

  11. NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing

    cs.AI 2026-03 reject novelty 6.0

    NeSy-Route is a 10,821-sample remote-sensing benchmark that generates constrained route-planning tasks from semantic masks plus A* search and evaluates MLLMs with a three-level symbolic protocol.

  12. Cryo-Bench: Benchmarking Foundation Models for Cryosphere Applications

    cs.CV 2026-03 conditional novelty 6.0

    A new benchmark shows a simple U-Net outperforms frozen geospatial foundation models on cryosphere segmentation, but fine-tuning with learning-rate tuning narrows or reverses the gap.

  13. MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training

    cs.CV 2026-02 conditional novelty 6.0

    MMEarth-Bench introduces five global multimodal environmental tasks, and TTT-MMR adapts pretrained models at test time by reconstructing 12 auxiliary modalities, improving average performance on both random and geogra...

  14. The View From Space: Navigating Instrumentation Differences with EOFMs

    cs.CV 2025-10 conditional novelty 6.0

    EOFM embeddings are strongly partitioned by sensor architecture, so matching spectral bands is not enough to make cross-sensor embedding search reliable.

  15. Treatment Geometry and Causal Identification with Earth Observation Data

    econ.EM 2026-07 accept novelty 5.0

    Defines treatment geometry and provides a protocol for making spatial and temporal exposure choices explicit in geospatial impact evaluations.

  16. Benchmarking Geospatial Foundation Models for Agriculture Applications

    cs.CV 2026-06 unverdicted novelty 5.0

    Benchmark of Prithvi, SpectralGPT, and SatMAE shows sharp performance drop under regional distribution shift, with models defaulting to common crops and missing rare ones.

  17. Location Is All You Need: Continuous Spatiotemporal Neural Representations of Earth Observation Data

    cs.CV 2026-04 unverdicted novelty 5.0

    LIANet encodes multi-temporal Earth observation data into a coordinate-based neural field that supports label-only fine-tuning for downstream tasks without access to raw imagery.

  18. How to Embed Matters: Evaluation of EO Embedding Design Choices

    cs.CV 2026-03 unverdicted novelty 5.0

    Transformer backbones with mean pooling and combined self-supervised embeddings yield robust, compact representations for EO tasks that are over 500x smaller than raw data.

  19. The Lov\'{a}sz Local Lemma: Foundations and Applications

    math.CO 2026-03 conditional novelty 5.0

    LEPA predicts geometrically transformed patch embeddings from context and transform parameters, lifting MRR from <0.2 (interpolation) to >0.8 while keeping competitive PANGAEA segmentation scores.

  20. SHRUG-FM: Reliability-Aware Foundation Models for Earth Observation

    cs.CV 2025-11 unverdicted novelty 5.0

    SHRUG-FM fuses geophysical OOD detection, embedding-space OOD detection, and predictive uncertainty via a shallow decision tree to let foundation models abstain from unreliable outputs on burn scar, flood, and landsli...

  21. Feature Extraction in the Remote Sensing Data Value Chain: A Systematic Review of Methods and Applications

    cs.CV 2025-10 unverdicted novelty 5.0

    A systematic review that introduces a framework for feature extraction in remote sensing, traces its evolution in the data value chain, and synthesizes trends toward unified representations and foundation models.

  22. Fusing Multi- and Hyperspectral Satellite Data for Harmful Algal Bloom Monitoring with Self-Supervised and Hierarchical Deep Learning

    cs.LG 2025-10 conditional novelty 5.0

    SIT-FUSE maps harmful algal bloom concentration and species from fused VIIRS/MODIS/Sentinel-3/PACE/TROPOMI data using self-supervised hierarchical clustering, with limited in-situ validation.

  23. Low-Rank Adaptation of Geospatial Foundation Models for Wildfire Mapping Using Sentinel-2 Data

    cs.CV 2026-05 unverdicted novelty 4.0

    LoRA-adapted Prithvi-v2 achieves the highest accuracy and best cross-domain generalization for burned-area mapping on Sentinel-2 data compared to full fine-tuning across 3,820 wildfire events.

  24. Scalable and Trustworthy Earth Observation Foundation Models

    cs.LG 2026-07 conditional novelty 3.0

    Remote-sensing foundation models need domain-specific design and evaluation around measurement physics and decision constraints; benchmark accuracy alone is insufficient for trustworthy EO deployment.

  25. The Lov\'{a}sz Local Lemma: Foundations and Applications

    math.CO 2026-03 unverdicted novelty 1.0

    An expository review presenting a pedagogically reformulated proof of the Lovász Local Lemma using unconditional inequalities, plus revisited applications and algorithmic perspectives.