{"id":"920a512a-5e85-44e8-812a-da010c455eea","arxiv_id":"2601.01558","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"AlphaEarth satellite embeddings modestly improve out-of-sample LSTM streamflow prediction and donor-basin selection over CAMELS attributes on 671 US basins.","lead":"This paper tests whether AlphaEarth satellite embeddings improve streamflow prediction in ungauged basins compared to traditional hand-crafted basin attributes. The authors find modest out-of-sample gains and better donor-basin selection using the embeddings, but a temporal mismatch between satellite data and test periods raises a caveat.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'no leakage' claim in §2.2 is not established: AEF embeddings are from 2017–2024 while the OOS test period is 2010–2014, so the 0.612-vs-0.553 NSE gain may reflect future land-surface information rather than superior transferability.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the temporal mismatch between AEF embeddings (2017–2024) and the CAMELS test period (2010–2014) and the unsupported 'no leakage' assertion. This is the most direct threat to the central claim because the OOS NSE difference is modest (0.059 median) and could plausibly be explained by future land-surface information. I considered other potential issues—the donor-selection experiment uses only five basins, the input dimensionalities differ (71 vs 24), and the cross-regime experiment in §5.2 shows AEF clusters generalize worse—but none of these directly undermine the main quantitative claim as much as the temporal-leakage confound. The paper is otherwise careful, provides code, and discusses the cross-regime trade-off, so a conditional verdict is appropriate: the claim should be accepted only after the temporal-window sensitivity is tested and the leakage assertion is either substantiated or removed.","tokens_in":15206,"tokens_out":5874,"duration_ms":72212,"concrete_test":"Re-run Experiment A's OOS protocol using only the 2017 annual AEF embedding (the earliest available) as the static descriptor, and separately using only the 2024 annual embedding. Compare the two resulting OOS NSE distributions to the reported 0.612 median and to the CAMELS-attribute median of 0.553. If the AEF advantage shrinks to <0.03 NSE or disappears with either single-year embedding, the 2017–2024 averaging is carrying post-test-period information and the 'no leakage' claim in §2.2 fails. If both single-year embeddings preserve the full advantage, the temporal-window sensitivity is not the driver, though a pre-2010 satellite-derived descriptor would be an even stronger confirmatory check.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central quantitative claim—AEF embeddings provide much stronger cross-regional generalization than CAMELS attributes—rests on the OOS median NSE gap (0.612 vs 0.553, §4.1). The paper's guard against leakage is the assertion in §2.2 that 'since the hydrological model training and testing phases predate the satellite observations, this setup inherently precludes any risk of data leakage.' This is not logically sufficient. The AEF embeddings are annual composites from 2017–2024, whereas the test period is 2010–2014; thus the input features for each test basin are derived from imagery acquired after the period being predicted. Land-surface states in 2017–2024 can be influenced by—and therefore encode information about—the same slow hydrological and ecological processes active in 2010–2014 (e.g., vegetation memory, groundwater trends, drought recovery, land-cover change). This is a form of target leakage: information from after the evaluation window enters the feature set. Because the CAMELS attributes are constructed from more historical data, the comparison is not clean: the AEF model may be exploiting future information, which would inflate its apparent cross-regional transferability. The paper acknowledges the temporal mismatch but does not test it; the 'no leakage' conclusion is assumed, not demonstrated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using AlphaEarth Foundation (AEF) satellite embeddings, 64-dimensional vectors derived from 2017–2024 satellite imagery, as static basin descriptors for large-sample hydrological modeling with the CAMELS-US dataset. It compares LSTM models trained with AEF embeddings against models trained with 17 traditional CAMELS physiographic attributes under in-sample (IS) and out-of-sample (OOS) conditions, reporting similar IS performance (median NSE ≈ 0.708 vs 0.706) but higher OOS median NSE (0.612 vs 0.553). The paper also evaluates similarity-based donor-basin selection for prediction in ungauged basins (PUB) using attribute similarity, an MLP fusion-embedding similarity, and AEF embedding similarity across five target basins, concluding that AEF similarity identifies more coherent donor sets and improves performance for small-to-moderate training-set sizes. Additional analyses examine mutual information between AEF embeddings and CAMELS attributes and a cross-regime generalization experiment based on clustering basins by each representation.","tokens_in":15561,"tokens_out":3932,"duration_ms":44461,"significance":"If the OOS result is unbiased, the paper would make a useful contribution by demonstrating that Earth Foundation Model embeddings can serve as transferable basin descriptors for PUB, with public embeddings and code availability. The mutual-information analysis and similarity-space visualization are constructive, and the paper raises an important question about the trade-off between discriminative precision and cross-regime generalization. However, the central quantitative claim is currently vulnerable to temporal information leakage, and the PUB experiment is based on only five subjectively selected basins. These issues must be resolved before the paper's main conclusions can be accepted.","major_comments":[{"comment":"The claim that the setup 'inherently precludes any risk of data leakage' is not established. AEF embeddings are annual composites from 2017–2024, while the OOS test period is 2010–2014. Land-surface conditions observed after the evaluation window can encode information about slow hydrological and ecological processes—vegetation memory, land-cover change, drought recovery, groundwater trends—that are correlated with streamflow in the test period. The CAMELS attributes, by contrast, are constructed from more historical data. The OOS median NSE gap (0.612 vs 0.553) may therefore reflect future information in the feature set rather than superior cross-regional transferability. This is load-bearing for the paper's central claim. Providing an explicit test or a carefully reasoned stability analysis is necessary; asserting temporal precedence of model training is not sufficient.","section":"§2.2, §4.1"},{"comment":"The statistical significance of the OOS comparison is overstated. The KS statistic for the OOS NSE distributions is only 0.0863, and the median difference is 0.059 NSE units. With 671 basins, a small KS statistic can nonetheless yield p ≈ 6×10⁻⁹, but this does not imply a practically large effect. The text repeatedly describes the AEF advantage as 'substantially' or 'much stronger,' which is not supported by the reported effect size. I recommend reporting bootstrap confidence intervals for the median difference (and for the KS statistic) and discussing the practical hydrological significance of a 0.06 median NSE gain.","section":"§4.1, Figure 4"},{"comment":"The PUB scaling experiment is based on five target basins that are described only as 'selected from distinct geographic and ecological regions' with no explicit selection criteria. This small, subjectively chosen sample cannot support general conclusions about AEF-based donor selection across the 671-basin dataset. The mean trend in Figure 8(f) has uncertainty bands, but no formal statistical test is performed across targets, and there is no correction for multiple comparisons across the 21 configurations per basin. I recommend either substantially expanding the target set, or reframing Experiment B as an exploratory case study rather than a general demonstration.","section":"§3, Experiment B; §4.2"},{"comment":"The cross-regime generalization experiment is confounded by the different number of clusters selected for each representation: 12 clusters for CAMELS attributes versus 9 for AEF embeddings. Leave-one-cluster-out therefore yields different training-set sizes and cluster compositions for the two models, so the observed difference in generalization cannot be attributed solely to the feature space. The text acknowledges this but does not control for it. A cleaner comparison would use the same number of clusters for both representations, or otherwise balance training-set size, to test whether the 'granularity-generalization trade-off' is a real property of AEF embeddings or an artifact of the clustering configuration.","section":"§5.2, Figure 11"}],"minor_comments":[{"comment":"The description of the IS implementation says 'the train period of 531 catchments,' but the paper states that all 671 CAMELS catchments were used. Please clarify whether this is a typo and, if 531 is intentional, explain the discrepancy.","section":"§3, Experiment A"},{"comment":"The training period ends in December 2004 and the test period begins in January 2010. The five-year gap is not discussed; please state whether this was intentional and whether it affects the comparison.","section":"§3, Experiment A"},{"comment":"The text says 'The results of Experiment B, summarized in Figure 7,' but the NSE scaling results appear in Figure 8. Please correct the cross-reference.","section":"§4.2, text near Figure 8"},{"comment":"The sentence 'since the hydrological model training and testing phases predate the satellite observations, this setup inherently precludes any risk of data leakage' should be reworded even if the authors decide to keep the temporal-mismatch assumption, because 'inherently precludes' is too strong given the mechanism described above.","section":"§2.2"},{"comment":"The MLP-embedding similarity is trained on the same streamflow prediction task, so it is a supervised, task-specific baseline rather than a task-agnostic representation. This should be stated explicitly when comparing it with AEF embeddings.","section":"§2.3, 3"}],"recommendation":"major_revision","confidential_remarks":"The central leakage concern is the main risk: if the OOS gain is driven by post-2014 land-surface information, the paper's headline result collapses. I would not accept in current form, but a revision that convincingly addresses the temporal mismatch and reframes the five-basin PUB experiment as exploratory could make this a publishable contribution. The availability of code and data is a positive feature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the first study I know of applying AlphaEarth embeddings to CAMELS PUB, and the main empirical claim—median OOS NSE 0.612 for AEF vs 0.553 for CAMELS attributes—is plausible but not as clean as the paper presents. The §2.2 statement that the setup 'inherently precludes any risk of data leakage' is incorrect as written, because the embeddings are annual composites from 2017–2024 while the test period is 2010–2014. That's a real temporal mismatch: land-surface states in the later window can encode slow hydrological processes active during the test window. It doesn't necessarily sink the paper, but it means the comparison is not apples-to-apples. The AEF model may be getting a boost from post-test information. The authors should either re-run with pre-2010 embeddings (if the AEF archive allows) or soften the claim and frame the result as a robustness check under covariate shift, not a leakage-free OOS evaluation.\n\nWhat the paper does well: it's careful in several ways. Five seeds, bootstrap resampling, KS tests, and a mutual-information analysis are all reasonable. The cluster-based cross-regime experiment is a nice idea and honestly reports a trade-off: AEF embeddings are sharper discriminators but generalize worse when entire clusters are withheld. The paper also ships code and data links, which I respect. Experiment B is transparently limited to five basins, and the finding that AEF similarity helps at small donor sizes but degrades with more heterogeneous donors is a useful empirical observation, even if the scope is narrow.\n\nSoft spots in proportion: the KS significance is inflated by the 671-basin sample, and the effect size is modest (about 0.06 median NSE). The MLP-embedding baseline is trained on the same streamflow task, so it's not a fully task-agnostic similarity definition, but the authors use it as a comparison and it performs worse, so that's not a circularity problem. The five target basins in Experiment B are subjectively chosen, and the paper says so.\n\nWho this is for: hydrologists working with large-sample benchmarks and anyone interested in using EO foundation model embeddings as static basin descriptors. It deserves a serious referee—the leak question is addressable and should drive the revision, but the core empirical contribution is worth engaging with.","headline":"A useful first benchmark for Earth foundation model embeddings in CAMELS PUB, but the §2.2 'no leakage' claim is wrong as stated and the central OOS gain is less clean than the paper suggests.","tokens_in":16033,"tokens_out":1800,"would_cite":true,"duration_ms":22488,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 64-dimensional satellite embedding of each basin, averaged over eight years of imagery, replaces hand-crafted catchment attributes and improves out-of-sample streamflow prediction.","keywords":["prediction in ungauged basins","Earth foundation models","satellite embeddings","static basin attributes","LSTM","CAMELS","basin similarity","cross-regional generalization"],"falsifier":"Recompute the 5-fold out-of-sample comparison using only satellite embeddings from 2017 (or from a foundation model trained on imagery ending before 2010). If the median NSE advantage over CAMELS attributes (0.612 vs 0.553) shrinks or reverses, the reported cross-regional generalization is largely an artifact of post-test land-surface information embedded in the averaged 2017-2024 vectors.","tokens_in":15103,"feed_emoji":"🛰️","tokens_out":4655,"duration_ms":45835,"temperature":0.7,"pith_summary":"This paper argues that basin descriptors learned by an Earth foundation model from satellite imagery are more transferable than the hand-crafted attributes traditionally used in large-sample hydrology. Using the 671-basin CAMELS benchmark, it finds that the two descriptor types perform about equally when predicting gauged basins in a new time period (median NSE 0.708 vs 0.706), but the satellite embeddings pull ahead when predicting basins never seen in training (0.612 vs 0.553). It also shows that choosing donor basins by cosine similarity in the embedding space yields more coherent, hydrologically consistent training sets, improving ungauged-basin prediction at small and moderate training sizes. If right, the result would give hydrologists a data-driven, task-agnostic way to characterize catchment similarity and would shift ungauged-basin prediction toward foundation-model representations.","feed_headline":"Satellite-learned descriptors beat handcrafted basin attributes","feed_subtitle":"In out-of-sample tests, median NSE rises from 0.553 to 0.612, and donor-basin selection becomes more coherent for ungauged prediction.","key_machinery":"The load-bearing object is the AlphaEarth Foundation embedding: a 64-dimensional vector per location, produced by a self-supervised model trained on global satellite imagery, averaged here over pixels and over the 2017-2024 period to give one static descriptor per basin. It replaces the 17 hand-crafted CAMELS attributes as the static input to an LSTM rainfall-runoff model, and its cosine similarity defines basin relatedness for donor-basin selection. The work it does is to supply a dense, integrated land-surface signal — vegetation, terrain, surface moisture, land cover — that the paper argues is more transferable and more discriminative than sparse expert attributes.","core_discovery":"On the paper's own terms, the central discovery is that AlphaEarth Foundation (AEF) embeddings — 64-dimensional vectors summarizing multi-year satellite imagery of each basin — constitute a more transferable static basin representation than the 17 hand-crafted CAMELS attributes. In 5-fold spatial cross-validation the AEF-based LSTM reaches a median out-of-sample NSE of 0.612 against 0.553 for the CAMELS-attribute model, while in-sample performance is nearly identical (0.708 vs 0.706). Mutual-information analysis shows partial overlap with terrain and vegetation attributes plus additional environmental signal not present in CAMELS. The embeddings also define a similarity space with high globa","pith_inferences":["Inference: the 2017-2024 embeddings overlap the 2010-2014 test period, so the claimed 'no leakage' is not guaranteed; the OOS gain could partly reflect land-surface changes observed after the test years, and re-running with pre-2014 embeddings would clarify this.","Inference: the compact 64-dim descriptor may serve as a general-purpose catchment signature beyond streamflow — e.g., for water-quality, drought, or ecohydrological prediction in ungauged locations — though the paper does not test these.","Inference: the granularity-generalization trade-off identified for whole-cluster withholding suggests a conditional or mixture-of-experts architecture, rather than a single shared network, as the natural next step to exploit AEF precision.","Inference: because AEF embeddings are derived globally from publicly available satellite data, the approach transfers to regions without CAMELS-style attributes, which is testable by applying it to non-US large-sample datasets."],"forward_implications":["If AEF embeddings generalize as claimed, regional LSTM models can drop hand-crafted basin attributes without losing in-sample skill and with gains in spatial out-of-sample skill.","Similarity-based donor selection in the embedding space can make ungauged-basin prediction more data-efficient, reaching high skill with only 100-300 similar basins.","The observed performance decline when donor sets grow too large implies that pooling all available basins is not automatically best, supporting similarity-aware training-set construction.","The cross-regime results imply that a single global model may underperform a segmented, similarity-aware set of models at current data scales."],"fun_headline_variants":["Satellite embeddings beat handcrafted attributes in ungauged basins","AI-learned basin descriptors from satellite images improve flow forecasts","Satellite embeddings sharpen ungauged basin predictions over manual attributes","Satellite-based basin embeddings lift ungauged streamflow prediction","Satellite embeddings guide donor-basin choice for ungauged prediction"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that a basin's 2017-2024 satellite embedding is a static description valid for the 1980-2014 modeling period; the paper's assertion that this setup 'inherently precludes' leakage is not logically forced, because the embeddings can contain land-surface information from after the test period.","fun_headline_variants_meta":{"raw":{"variants":["Satellite embeddings beat handcrafted attributes in ungauged basins","AI-learned basin descriptors from satellite images improve flow forecasts","Satellite embeddings sharpen ungauged basin predictions over manual attributes","Satellite-based basin embeddings lift ungauged streamflow prediction","Satellite embeddings guide donor-basin choice for ungauged prediction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000726,"raw_usage":{"total_tokens":3078,"prompt_tokens":719,"completion_tokens":2359,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":463,"completion_tokens_details":{"reasoning_tokens":2278}},"tokens_in":463,"tokens_out":2359,"duration_ms":17094,"temperature":1.0,"reasoning_tokens":2278,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T12:45:33.276086+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the 5-fold out-of-sample comparison using only satellite embeddings from 2017 (or from a foundation model trained on imagery ending before 2010). If the median NSE advantage over CAMELS attributes (0.612 vs 0.553) shrinks or reverses, the reported cross-regional generalization is largely an artifact of post-test land-surface information embedded in the averaged 2017-2024 vectors.","supporting_citations":[],"review_version":1}