REVIEW 25 cited by
PANGAEA: A Global and Inclusive Benchmark for Geospatial Foundation Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
PANGAEA: A Global and Inclusive Benchmark for Geospatial Foundation Models
read the original abstract
Geospatial Foundation Models (GFMs) have emerged as powerful tools for extracting representations from Earth observation data, but their evaluation remains inconsistent and narrow. Existing works often evaluate on suboptimal downstream datasets and tasks, that are often too easy or too narrow, limiting the usefulness of the evaluations to assess the real-world applicability of GFMs. Additionally, there is a distinct lack of diversity in current evaluation protocols, which fail to account for the multiplicity of image resolutions, sensor types, and temporalities, which further complicates the assessment of GFM performance. In particular, most existing benchmarks are geographically biased towards North America and Europe, questioning the global applicability of GFMs. To overcome these challenges, we introduce PANGAEA, a standardized evaluation protocol that covers a diverse set of datasets, tasks, resolutions, sensor modalities, and temporalities. It establishes a robust and widely applicable benchmark for GFMs. We evaluate the most popular GFMs openly available on this benchmark and analyze their performance across several domains. In particular, we compare these models to supervised baselines (e.g. UNet and vanilla ViT), and assess their effectiveness when faced with limited labeled data. Our findings highlight the limitations of GFMs, under different scenarios, showing that they do not consistently outperform supervised models. PANGAEA is designed to be highly extensible, allowing for the seamless inclusion of new datasets, models, and tasks in future research. By releasing the evaluation code and benchmark, we aim to enable other researchers to replicate our experiments and build upon our work, fostering a more principled evaluation protocol for large pre-trained geospatial models. The code is available at https://github.com/VMarsocci/pangaea-bench.
Forward citations
Cited by 25 Pith papers
-
Moonstone: A Multimodal Foundation Model and Benchmark for Lunar Remote Sensing
MG-MAE pretrained on a new 28-channel lunar map outperforms ImageNet, vanilla MAE, and EO foundation models across six classification, regression, and segmentation tasks.
-
Uncertainty-aware tree height change regression
Introduces the CHC dataset of 3 m continuous canopy height differences with uncertainties and the uncertainty-aware change regression task for fine-tuning GFMs on PlanetScope time series.
-
UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observation
UniverSat is a ViT-style model with a universal patch encoder enabling self-supervised training on heterogeneous multimodal Earth observation data from varying resolutions and sensors.
-
GeoNatureAgent Benchmark: Benchmarking LLM Agents for Environmental Geospatial Analysis Across Frontier and Open-Weight Foundation Models
GeoNatureAgent Benchmark tests seven LLMs on 93 tasks via a production geospatial API, with Claude Sonnet 4 at 60.8% and DeepSeek V3.2 offering near performance at 11x lower cost while all models fail on close-value c...
-
SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals
SDGBiasBench reveals intrinsic SDG biases in VLMs driven by priors rather than evidence, and CADE mitigates them with up to 25% accuracy gains and 12-point MAE reductions.
-
Does Your Wildfire Prediction Model Actually Work, or Just Score Well?
WILDFIRE-FM is the first wildfire-specific Earth foundation model, paired with a fixed-contract evaluation framework that demonstrates wildfire model transfer conclusions depend strongly on evaluation design and task ...
-
NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing
NeSy-Route supplies 10,821 optimally labeled remote-sensing route-planning tasks plus a three-level neuro-symbolic protocol that reveals major perception and planning deficits in current MLLMs.
-
Does Your Wildfire Prediction Model Actually Work, or Just Score Well?
Introduces WILDFIRE-FM and a fixed-contract evaluation framework demonstrating that wildfire model transfer conclusions depend strongly on evaluation design and task formulation.
-
No One Knows the State of the Art in Geospatial Foundation Models
An audit of 152 papers reveals that geospatial foundation models lack standardized evaluations, training controls, and weight releases, so no one knows the state of the art.
-
LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation
LithoBench is a new multi-level benchmark showing that existing large multimodal models have substantial limitations in geological semantic understanding for remote sensing lithology interpretation.
-
NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing
NeSy-Route is a 10,821-sample remote-sensing benchmark that generates constrained route-planning tasks from semantic masks plus A* search and evaluates MLLMs with a three-level symbolic protocol.
-
Cryo-Bench: Benchmarking Foundation Models for Cryosphere Applications
A new benchmark shows a simple U-Net outperforms frozen geospatial foundation models on cryosphere segmentation, but fine-tuning with learning-rate tuning narrows or reverses the gap.
-
MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training
MMEarth-Bench introduces five global multimodal environmental tasks, and TTT-MMR adapts pretrained models at test time by reconstructing 12 auxiliary modalities, improving average performance on both random and geogra...
-
The View From Space: Navigating Instrumentation Differences with EOFMs
EOFM embeddings are strongly partitioned by sensor architecture, so matching spectral bands is not enough to make cross-sensor embedding search reliable.
-
Treatment Geometry and Causal Identification with Earth Observation Data
Defines treatment geometry and provides a protocol for making spatial and temporal exposure choices explicit in geospatial impact evaluations.
-
Benchmarking Geospatial Foundation Models for Agriculture Applications
Benchmark of Prithvi, SpectralGPT, and SatMAE shows sharp performance drop under regional distribution shift, with models defaulting to common crops and missing rare ones.
-
Location Is All You Need: Continuous Spatiotemporal Neural Representations of Earth Observation Data
LIANet encodes multi-temporal Earth observation data into a coordinate-based neural field that supports label-only fine-tuning for downstream tasks without access to raw imagery.
-
How to Embed Matters: Evaluation of EO Embedding Design Choices
Transformer backbones with mean pooling and combined self-supervised embeddings yield robust, compact representations for EO tasks that are over 500x smaller than raw data.
-
The Lov\'{a}sz Local Lemma: Foundations and Applications
LEPA predicts geometrically transformed patch embeddings from context and transform parameters, lifting MRR from <0.2 (interpolation) to >0.8 while keeping competitive PANGAEA segmentation scores.
-
SHRUG-FM: Reliability-Aware Foundation Models for Earth Observation
SHRUG-FM fuses geophysical OOD detection, embedding-space OOD detection, and predictive uncertainty via a shallow decision tree to let foundation models abstain from unreliable outputs on burn scar, flood, and landsli...
-
Feature Extraction in the Remote Sensing Data Value Chain: A Systematic Review of Methods and Applications
A systematic review that introduces a framework for feature extraction in remote sensing, traces its evolution in the data value chain, and synthesizes trends toward unified representations and foundation models.
-
Fusing Multi- and Hyperspectral Satellite Data for Harmful Algal Bloom Monitoring with Self-Supervised and Hierarchical Deep Learning
SIT-FUSE maps harmful algal bloom concentration and species from fused VIIRS/MODIS/Sentinel-3/PACE/TROPOMI data using self-supervised hierarchical clustering, with limited in-situ validation.
-
Low-Rank Adaptation of Geospatial Foundation Models for Wildfire Mapping Using Sentinel-2 Data
LoRA-adapted Prithvi-v2 achieves the highest accuracy and best cross-domain generalization for burned-area mapping on Sentinel-2 data compared to full fine-tuning across 3,820 wildfire events.
-
Scalable and Trustworthy Earth Observation Foundation Models
Remote-sensing foundation models need domain-specific design and evaluation around measurement physics and decision constraints; benchmark accuracy alone is insufficient for trustworthy EO deployment.
-
The Lov\'{a}sz Local Lemma: Foundations and Applications
An expository review presenting a pedagogically reformulated proof of the Lovász Local Lemma using unconditional inequalities, plus revisited applications and algorithmic perspectives.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.