REVIEW 7 cited by
SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Geographic information is essential for modeling tasks in fields ranging from ecology to epidemiology. However, extracting relevant location characteristics for a given task can be challenging, often requiring expensive data fusion or distillation from massive global imagery datasets. To address this challenge, we introduce Satellite Contrastive Location-Image Pretraining (SatCLIP). This global, general-purpose geographic location encoder learns an implicit representation of locations by matching CNN and ViT inferred visual patterns of openly available satellite imagery with their geographic coordinates. The resulting SatCLIP location encoder efficiently summarizes the characteristics of any given location for convenient use in downstream tasks. In our experiments, we use SatCLIP embeddings to improve prediction performance on nine diverse location-dependent tasks including temperature prediction, animal recognition, and population density estimation. Across tasks, SatCLIP consistently outperforms alternative location encoders and improves geographic generalization by encoding visual similarities of spatially distant environments. These results demonstrate the potential of vision-location models to learn meaningful representations of our planet from the vast, varied, and largely untapped modalities of geospatial data.
Forward citations
Cited by 7 Pith papers
-
RainShift: A Benchmark for Precipitation Downscaling Across Geographies
RainShift is a global benchmark showing that precipitation downscaling models lose up to 30% accuracy when applied to unseen Global South regions, and input quantile mapping recovers some of that loss.
-
CanadaFireSat: Toward high-resolution wildfire forecasting with multiple modalities
Introduces a multi-modal 100m wildfire forecasting benchmark for Canada and shows deep learning models benefit from fusing Sentinel-2 imagery with environmental predictors.
-
Trees as Gaussians: Large-Scale Individual Tree Mapping
A deep learning system detects individual large trees globally in 3 m PlanetScope imagery by regressing Gaussian heatmaps trained on 14 billion lidar-derived pseudo-labels.
-
GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space
GT-Loc jointly predicts capture location, hour, and month via retrieval in a shared image/location/time embedding space, using a toroidal temporal metric loss.
-
GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields
GeoProg3D combines a georeferenced hierarchical 3D language field, geographic vision APIs, and LLM-generated programs to answer natural-language queries about city-scale 3D scenes, and includes a new 952-query benchma...
-
Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
Vision Transformers with SatCLIP location embeddings beat prior CNN baselines on thaw slump and ice-wedge polygon detection test splits, but not on infrastructure detection.
-
Enriching Location Representation with Detailed Semantic Information
Adding POI names to multimodal contrastive location embeddings improves land use classification, socioeconomic mapping, and location retrieval over type-only and traditional baselines.
Discussion (0). Continue with ORCID to comment.