{"id":"b4ce221c-248a-4b21-95f6-4bc1812ae442","arxiv_id":"2501.12535","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Balanced, globally representative pre-training data generally outperforms region-specific sampling for two geospatial foundation models in few-shot downstream tasks, and the advantage shrinks as finetuning data grows.","lead":"This paper tests whether the geographic mix of satellite images used to pre-train geospatial AI models changes how well the models handle agriculture and ecosystem mapping tasks. It finds that globally balanced data usually beats region-specific data in few-shot settings, but the best choice depends on the model architecture.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'balanced beats clustered' claim is contradicted by the paper's own Table 1 on Presto/CropHarvest: World Cities matches UAR on all continents, so the headline rests on a single model/task pair.","rationale":"Good-faith reading: the paper asks a fair question, equalizes sample counts, uses 50 seeds, and disaggregates by continent, which is more careful than much of the existing GFM literature. The concern is not about fabrication or about the raw numbers; it is about what the numbers license. Table 1 itself shows that the only composition consistently below UAR is Natural Forest, and that World Cities, the other 'clustered' condition, is at parity for Presto/CropHarvest. Since Presto/CropHarvest is the sole pixel-timeseries evidence, the claim that balanced sampling 'generally' outperforms clustered sampling is one experiment away from being true. A paired significance analysis is the cheapest decisive check: the seeds already exist, the differences are small, and standard errors around 0.02–0.03 across 50 runs imply that paired tests would be much more sensitive than the displayed marginal means. Depending on the outcome, the verdict stays CONDITIONAL; the condition should be that the general claim holds only after paired tests and only on the tasks and architectures where separation is reproducible. This does not change the reader's CONDITIONAL verdict but pinpoints the load-bearing evidence to check.","tokens_in":14040,"tokens_out":6725,"duration_ms":73481,"concrete_test":"Recompute Table 1 from the stored per-seed results: for each of the 50 finetuning seeds, form paired continent-wise F1 differences (UAR minus World Cities, UAR minus Natural Forest) for CropHarvest with Random Forest, aggregate across continents, and run a paired permutation test with per-composition correction. If the UAR-vs-World-Cities effect is not significant (or the effect size is below 0.01), the central claim must be restricted to EcoRegions/SatCLIP; apply the same paired analysis to the Appendix D finetuning heads.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weakness is internal to Section 4 and the Conclusion. The paper's central assertion (abstract, §4, Conclusion) is that balanced sampling techniques 'usually outperform' clustered or region-specific compositions. But Table 1, using the authors' own highlight threshold (scores at least 2% below UAR), shows World Cities with no highlighted entries for CropHarvest on any continent; in Africa and North America it is numerically above UAR, and elsewhere it is within 0.01–0.02. Natural Forest is clearly worse in several continents, but World Cities behaves like a balanced composition for Presto. The only large, consistent separation appears in SatCLIP/EcoRegions, where the task labels are biome classes and the balanced pre-training sets were deliberately stratified over biomes, so the advantage is partly built into the task design. Because no paired significance tests are reported (only mean ± std over 50 seeds), we cannot tell whether the small CropHarvest differences are real or noise. Therefore the global recommendation 'balanced and global representative sampling generally outperform clustered' is not supported by the Presto experiment; the evidence supports at most a task- and architecture-specific conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how the spatial distribution of pre-training data affects the downstream performance of geospatial foundation models. For two GFMs (Presto and SatCLIP), the authors construct five pre-training compositions (uniform random, stratified by continent, stratified by biome, natural-forest-only, and world-cities-only), pre-train each model on equal-sized subsets, and evaluate continent-wise few-shot performance on CropHarvest and EcoRegions with 50 random seeds. The paper claims that balanced, globally representative sampling techniques outperform clustered or region-specific compositions, and that differences diminish as finetuning data grows.","tokens_in":14245,"tokens_out":2732,"duration_ms":28281,"significance":"If established, the paper would provide concrete guidance for GFM pre-training data curation, a topic that is relatively underexplored compared to architecture and pretext-task design. The study is systematic in several respects: it uses two distinct model families, multiple finetuning classifiers, repeated seeds, and a no-pre-training baseline. The authors also transparently acknowledge scope limitations in Appendix E. However, the central claim is only partially supported by the reported results, and the evidence requires either statistical strengthening or a more careful qualification.","major_comments":[{"comment":"The claim that 'all balanced data sampling techniques (i.e., UAR, stratified continent/biome) outperform clustered techniques' is contradicted by the Presto/CropHarvest results. World Cities, the clustered population-centric composition, has no entries highlighted as at least 2% below UAR, and it is numerically above UAR for Africa (0.72 vs 0.71) and North America (0.81 vs 0.80). The only clustered composition that is consistently worse is Natural Forest, and even that equivalence fails for South America and Oceania. The conclusion in Section 5 ('balanced and global representative sampling techniques generally outperform clustered or region-specific compositions') overstates what Table 1 shows for Presto; the assertion needs to be restricted to Natural Forest or to SatCLIP/EcoRegions, or supported with additional evidence that World Cities is meaningfully worse.","section":"Section 4, Table 1"},{"comment":"No significance testing is reported. For the Presto/CropHarvest comparisons, the largest difference between UAR and World Cities on any continent is 0.02, while the standard deviations are 0.02-0.03 across 50 runs. Without paired significance tests (e.g., a paired bootstrap or permutation test on the 50 seeds), the statement that balanced techniques 'usually outperform' clustered techniques is not statistically grounded. Please report confidence intervals for the pairwise differences or explicit significance tests, at least for the headline Table 1 comparisons.","section":"Section 4, Table 1"},{"comment":"The EcoRegions task is a 14-class biome classification, and the Stratified Biome pre-training composition is explicitly built by sampling equal numbers per biome using the same Dinerstein et al. (2017) scheme that defines the task labels. This creates a direct alignment between pre-training distribution and downstream label structure, which may explain the large SatCLIP advantage of balanced over clustered compositions. The paper does not acknowledge this potential confound when it generalizes to 'globally diverse pre-training data' being beneficial. The conclusion should be qualified to note that part of the observed effect may be task-specific label alignment rather than generic geographic balance.","section":"Section 3.2 and Section 4, Table 1"},{"comment":"For Presto, all five compositions are resampled from a single existing pool of approximately 22 million samples (Tseng et al. 2023). The paper does not report the spatial or environmental distribution of this pool. If the pool already has near-global coverage or is biased toward certain regions, the 'Natural Forest' and 'World Cities' subsets may not be as distinct from UAR as intended, weakening the manipulation that the central comparison depends on. Please report the pool's continent/biome distribution, or at least a measure of overlap between the compositions, to verify that the sampling strategies instantiate different intended distributions.","section":"Appendix B, Section 3.3"}],"minor_comments":[{"comment":"There is a typo in the abstract: 'thegeographic' should be 'the geographic'.","section":"Abstract"},{"comment":"The reference for Manas et al. (2021) contains 'uUncurated' which appears to be a typo for 'Uncurated'.","section":"References"},{"comment":"There are several spelling errors such as 'fineuning' and 'finetuning' used inconsistently; please standardize the spelling.","section":"Appendix C"},{"comment":"The caption for Figure 3 lists all six composition names as subcaptions for every subplot, which is confusing; consider labeling each subplot directly with its composition name.","section":"Figure 3"},{"comment":"The sentence 'The reason behind this discrepancy could be the architectural design of the model' is presented without supporting analysis; please mark it explicitly as a hypothesis and, if possible, support it with an analysis of learned representations or a toy experiment.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The paper shares authors with the Presto model (Tseng et al. 2023) and SatCLIP (Klemmer et al. 2023), and the Presto pre-training pool is taken directly from that prior work. This is not a reason to reject, but the authors should be encouraged to disclose the pool's creation procedure and any potential bias it introduces, as the composition manipulations for Presto depend on the pool's internal distribution. Also, the paper's framing as 'first study' in the introduction is somewhat strong given related work in general CV cited by the authors themselves; the contribution is better framed as the first systematic comparison for GFMs."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is a useful first cut at a real question—how the spatial distribution of pre-training data affects geospatial foundation models—but its headline claim is stronger than the evidence. The cleanest result is SatCLIP/EcoRegions, where balanced sampling clearly beats both clustered strategies. On Presto/CropHarvest, World Cities performs at parity with balanced sampling on every continent, so the \"balanced beats clustered\" rule fails for that model-task pair. The paper's own Section 4 acknowledges this, but the abstract and conclusion still say balanced techniques \"usually outperform\" clustered ones; that \"usually\" is doing a lot of work.\n\nWhat's actually new: this is the first systematic comparison of five sampling strategies for GFMs, with continent-wise few-shot evaluation, 50 seeds per setting, and a zero-pretraining baseline. That is a sound experimental skeleton. The architecture-dependent result—World Cities is fine for a pixel-timeseries model like Presto but poor for a location encoder like SatCLIP—is a genuinely interesting observation worth following up.\n\nThe soft spots are real but not fatal. No significance tests are reported, and several CropHarvest differences are 1–3 F1 points with overlapping standard deviations; we can't distinguish signal from noise there. The EcoRegions advantage is partly built into the task design, since the labels are biomes and two of the balanced compositions are stratified over biomes. The Presto pool is resampled from the original Presto dataset, which shares authors; the paper doesn't characterize the pool's internal spatial distribution, so the intended contrast between compositions may be weaker than advertised. No code or data is released.\n\nThe limitations section is candid and lists sensible next steps, which earns credit. The paper is worth a serious referee—the question is important and the design is mostly there—but it needs a revision: add paired significance tests (or at least effect sizes), temper the abstract to match the actual results, and ideally release the data and code. I'd cite it as the first comparison of its kind, but I wouldn't treat the guidance as settled.","headline":"Useful first systematic comparison of pre-training data sampling for GFMs, but the central 'balanced beats clustered' claim rests mostly on one model-task pair and needs significance testing and tempering.","tokens_in":14795,"tokens_out":2535,"would_cite":true,"duration_ms":24675,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Balanced and globally representative pre-training data generally outperform region-clustered data for geospatial foundation models in few-shot settings.","keywords":["geospatial foundation models","pre-training data distribution","data sampling strategies","few-shot learning","satellite imagery","Earth observation","representation learning","remote sensing"],"falsifier":"Compute the land-cover and continental histograms of the ~22-million-sample Presto pool and the overlap between its 'uniform at random' and 'natural forest' subsets; if the two compositions are nearly indistinguishable in their spatial or land-cover statistics, the central comparison collapses. A stronger test is to pre-train each model on pools built from scratch per composition, as the paper did for SatCLIP, and check whether the balanced-vs-clustered ranking persists.","tokens_in":13811,"feed_emoji":"🌍","tokens_out":5555,"duration_ms":50687,"temperature":0.7,"pith_summary":"This paper asks whether the geographic distribution of pre-training data changes how well geospatial foundation models perform on downstream Earth-observation tasks. It pre-trains two models — a pixel-timeseries model and a location-encoder model — on five equal-size data compositions ranging from globally balanced to regionally clustered, then fine-tunes each on continent-specific subsets of two globally distributed tasks using only 100 labeled samples per continent. The paper's central finding is that balanced and globally representative compositions (uniform at random, stratified by continent, stratified by biome) generally outperform clustered compositions (intact forests only, or a 50-km radius around the world's largest cities), and that all pretrained variants beat a no-pretraining baseline. It also finds that the performance gaps between compositions shrink as the number of fine-tuning samples grows, so the choice of sampling strategy matters most in the few-shot regime. If correct, the results give GFM developers a concrete, low-cost way to improve model quality: spend the sampling budget on guaranteeing global coverage rather than on simply amassing more data.","feed_headline":"Balanced global data beats clustered pretraining for geospatial AI","feed_subtitle":"Uniform or continent-stratified pre-training data beats forest-only and city-only samples in few-shot tests.","key_machinery":"The load-bearing instrument is a controlled resampling pipeline. From a fixed global pool (the ~22 million-sample Presto pool, or Sentinel-2 patches retrieved via a cloud catalog for SatCLIP), the authors create five pre-training compositions of equal size: uniform-at-random over land, stratified by continent, stratified by biome, all within intact-forest cover, and all within 50 km of the world's 10,000 most populated cities. Each composition is used to pre-train one pixel-timeseries GFM and one location-encoder GFM, and each pretrained model is then fine-tuned continent-wise on two global downstream tasks (crop vs. non-crop classification and eco-region classification) with 100 samples per continent, repeated 50 times to average over sampling noise. The pipeline isolates the spatial distribution of pre-training data as the independent variable while holding model architecture, pre-training configuration, pre-training data volume, and fine-tuning procedure fixed.","core_discovery":"The paper claims that, for two structurally different geospatial foundation models, the spatial composition of pre-training data is a first-order factor in downstream few-shot performance. Specifically, all balanced sampling techniques (uniform-at-random, continent-stratified, and biome-stratified) yield approximately equal F1 scores, and these balanced compositions match or exceed clustered compositions (Natural Forest and World Cities) across six continents on both the CropHarvest and EcoRegions tasks. The authors further claim that the relative ranking of compositions is not universal: the city-clustered composition performs on par with balanced ones for the pixel-timeseries model Presto but poorly for the location-encoder model SatCLIP, which they attribute to architectural differences in how each model uses location information.","pith_inferences":["If the generality holds, a practical corollary follows that GFM teams should measure the spatial coverage of their pre-training pool before scaling up data collection, since a biased pool cannot be rescued simply by resampling balanced subsets from it.","The architecture-dependence result suggests a testable mechanism: location encoders may internalize the spatial prior directly, so removing location input or adding positional augmentation might reduce the gap between balanced and clustered pre-training.","An implicit extension is that deliberately region-specific pre-training could be optimal for region-specific downstream tasks; the paper's continent-wise results provide a template for testing this trade-off at finer granularity than continents."],"forward_implications":["GFM pre-training datasets should be curated to guarantee global coverage across continents and biomes rather than clustered by region or environment, at least for few-shot global downstream tasks.","Among balanced sampling strategies, the exact method matters little: uniform-at-random, continent-stratified, and biome-stratified perform about equally.","Clustered sampling is not categorically harmful; it can match balanced sampling on continents where the cluster naturally dominates (e.g., Natural Forest for crop classification in South America and Oceania).","The benefits of any particular sampling strategy fade as fine-tuning data grows, so sampling decisions should be weighted most heavily in low-label deployment regimes.","Architecture interacts with data distribution: a location-encoder model is more sensitive to spatially clustered pre-training data than a pixel-timeseries model."],"supporting_citations":[{"why":"Supplies the Presto architecture, its pre-training configuration, and the ~22-million-sample global pool from which all Presto data compositions were resampled.","marker":"Tseng et al. 2023"},{"why":"Supplies the SatCLIP location-encoder model, its training configuration, and the uniform-at-random global sampling approach adopted for the UAR composition.","marker":"Klemmer et al. 2023"},{"why":"Defines the biome/ecoregion strata used for the Stratified Biome composition and provides the EcoRegions downstream task.","marker":"Dinerstein et al. 2017"},{"why":"Defines the global intact forest cover used to construct the Natural Forest clustered composition.","marker":"Potapov et al. 2017"},{"why":"The source of the city-based sampling approach (50 km around 10,000 cities) that the World Cities composition emulates.","marker":"Manas et al. 2021"},{"why":"Provides the CropHarvest dataset used as the downstream task for Presto.","marker":"Tseng et al. 2021"},{"why":"Previous finding that pre-training distribution effects diminish with more fine-tuning data, which the paper's convergence results corroborate.","marker":"Entezari et al. 2023"},{"why":"Motivates the focus on data subgroup composition and quality rather than sheer quantity in ML training.","marker":"Rolf et al. 2021"}],"fun_headline_variants":["Balanced pre-training data boosts geospatial AI performance","Global data diversity wins for geospatial foundation models","Spatial data mix matters more than architecture for GFMs","Uniform pre-training data beats location-clustered for GFMs","Pre-training data spread shapes geospatial model accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that the existing data pools can actually instantiate the intended spatial distributions — in particular, that resampling the ~22-million-sample Presto pool yields compositions as distinct as 'natural forest' and 'uniform at random' are meant to be; if the pool is already skewed geographically or by land cover, the observed differences cannot be attributed to sampling strategy.","fun_headline_variants_meta":{"raw":{"variants":["Balanced pre-training data boosts geospatial AI performance","Global data diversity wins for geospatial foundation models","Spatial data mix matters more than architecture for GFMs","Uniform pre-training data beats location-clustered for GFMs","Pre-training data spread shapes geospatial model accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000188,"raw_usage":{"total_tokens":1298,"prompt_tokens":878,"completion_tokens":420,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":342}},"tokens_in":494,"tokens_out":420,"duration_ms":4492,"temperature":1.0,"reasoning_tokens":342,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:05:48.213311+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the land-cover and continental histograms of the ~22-million-sample Presto pool and the overlap between its 'uniform at random' and 'natural forest' subsets; if the two compositions are nearly indistinguishable in their spatial or land-cover statistics, the central comparison collapses. A stronger test is to pre-train each model on pools built from scratch per composition, as the paper did for SatCLIP, and check whether the balanced-vs-clustered ranking persists.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the biome/ecoregion strata used for the Stratified Biome composition and provides the EcoRegions downstream task."},{"cited_title":"C.; Laestadius, L.; Turubanova, S.; Yaroshenko, A.; Thies, C.; Smith, W.; Zhuravleva, I.; Komarova, A.; Minnemeyer, S.; et al","cited_arxiv_id":null,"evidence_quote":"Defines the global intact forest cover used to construct the Natural Forest clustered composition."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The source of the city-based sampling approach (50 km around 10,000 cities) that the World Cities composition emulates."},{"cited_title":"L.; and Kerner, H","cited_arxiv_id":null,"evidence_quote":"Provides the CropHarvest dataset used as the downstream task for Presto."},{"cited_title":"The Role of Pre-training Data in Transfer Learning","cited_arxiv_id":"2302.13602","evidence_quote":"Previous finding that pre-training distribution effects diminish with more fine-tuning data, which the paper's convergence results corroborate."},{"cited_title":"T.; Recht, B.; and Jordan, M","cited_arxiv_id":null,"evidence_quote":"Motivates the focus on data subgroup composition and quality rather than sheer quantity in ML training."}],"review_version":1}