Pith. sign in

REVIEW 7 cited by

SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.17179 v3 pith:W4ONWYDW submitted 2023-11-28 cs.CV cs.AIcs.CYcs.LG

classification cs.CVcs.AIcs.CYcs.LG
keywords locationsatclipgeographictasksglobalimagerysatellitecharacteristics
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Geographic information is essential for modeling tasks in fields ranging from ecology to epidemiology. However, extracting relevant location characteristics for a given task can be challenging, often requiring expensive data fusion or distillation from massive global imagery datasets. To address this challenge, we introduce Satellite Contrastive Location-Image Pretraining (SatCLIP). This global, general-purpose geographic location encoder learns an implicit representation of locations by matching CNN and ViT inferred visual patterns of openly available satellite imagery with their geographic coordinates. The resulting SatCLIP location encoder efficiently summarizes the characteristics of any given location for convenient use in downstream tasks. In our experiments, we use SatCLIP embeddings to improve prediction performance on nine diverse location-dependent tasks including temperature prediction, animal recognition, and population density estimation. Across tasks, SatCLIP consistently outperforms alternative location encoders and improves geographic generalization by encoding visual similarities of spatially distant environments. These results demonstrate the potential of vision-location models to learn meaningful representations of our planet from the vast, varied, and largely untapped modalities of geospatial data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 19 citations worldwide. Full citation record

  1. RainShift: A Benchmark for Precipitation Downscaling Across Geographies

    cs.CV 2025-07 conditional novelty 7.0 of 10

    RainShift is a global benchmark showing that precipitation downscaling models lose up to 30% accuracy when applied to unseen Global South regions, and input quantile mapping recovers some of that loss.

  2. CanadaFireSat: Toward high-resolution wildfire forecasting with multiple modalities

    cs.CV 2025-06 conditional novelty 6.5 of 10

    Introduces a multi-modal 100m wildfire forecasting benchmark for Canada and shows deep learning models benefit from fusing Sentinel-2 imagery with environmental predictors.

  3. Trees as Gaussians: Large-Scale Individual Tree Mapping

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A deep learning system detects individual large trees globally in 3 m PlanetScope imagery by regressing Gaussian heatmaps trained on 14 billion lidar-derived pseudo-labels.

  4. GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space

    cs.CV 2025-07 conditional novelty 6.0 of 10

    GT-Loc jointly predicts capture location, hour, and month via retrieval in a shared image/location/time embedding space, using a toroidal temporal metric loss.

  5. GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GeoProg3D combines a georeferenced hierarchical 3D language field, geographic vision APIs, and LLM-generated programs to answer natural-language queries about city-scale 3D scenes, and includes a new 952-query benchma...

  6. Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Vision Transformers with SatCLIP location embeddings beat prior CNN baselines on thaw slump and ice-wedge polygon detection test splits, but not on infrastructure detection.

  7. Enriching Location Representation with Detailed Semantic Information

    cs.CE 2025-06 conditional novelty 5.0 of 10

    Adding POI names to multimodal contrastive location embeddings improves land use classification, socioeconomic mapping, and location retrieval over type-only and traditional baselines.

Pith tools