{"id":"cc3b25d0-7f2c-4dc4-8c86-9dffd31dfe99","arxiv_id":"2608.04393","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"For factor-derived proxy targets, keeping the degraded formula estimate as an anchored prediction and learning a gated correction (RASPL) improves degradation and tail robustness over direct prediction and over treating the formula as a feature.","lead":"This paper studies machine learning models trained on proxy targets computed from known scientific factors, such as RUSLE soil-loss estimates. It introduces diagnostics to distinguish true robustness from equation reconstruction, and a residual-learning framework (RASPL) that anchors predictions to the degraded formula estimate.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline MLP direct baseline is information-deprived (no R, LS, or C), so the claimed RASPL-vs-direct gain may be an information-access artifact rather than evidence for formula preservation.","rationale":"The reader's weakest assumption concerned external validity: the proxy target is generated by the same RUSLE equation used as the anchor, and no independent soil-loss measurements are used. That is a legitimate limitation. However, the more immediate and load-bearing concern is internal to the paper's headline comparison. The matched direct MLP baseline is not matched in information: Direct-MLP-STATS receives only the reliability vector (K statistics and degradation indicators), while RASPL and the formula-feature baseline receive the full product through the formula estimate. A direct predictor without R, LS, or C cannot possibly reconstruct A, so its low R2 is an artifact of missing inputs, not a demonstration that direct prediction is weak. The proposed check directly tests whether giving the direct baseline the same factor information as the formula branch removes the claimed advantage. If it does, the central claim's support is substantially weakened; if not, the concern is resolved. This does not fully address the reader's external-validity concern, but it is more fundamental because it affects the internal validity of the main experimental contrast. The verdict should remain CONDITIONAL, with an explicit added condition: rerun the matched comparisons with information-equivalent direct baselines before the formula-preservation principle is accepted.","tokens_in":11970,"tokens_out":10848,"duration_ms":143112,"concrete_test":"Re-run the Table 2 protocol with Direct-MLP-STATS augmented by the four center-pixel channel values (R, LS, C, and degraded K̃) in addition to the existing reliability vector, keeping all other choices identical (seeds, training, checkpoint selection, evaluation). If the augmented direct model's macro R2_all over the three degraded regimes approaches or exceeds RASPL-MLP-STATS's 0.834, the claimed formula-preservation advantage is an information-access artifact. If it remains near 0.10, the concern is refuted and the MLP comparison is valid as reported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central comparison in Table 2 is not information-matched. Section 'Matched Direct and Formula-Feature Baselines' says direct baselines retain the same contextual encoders and regime-matched inputs as their RASPL counterparts. For RASPL-MLP-STATS, the contextual encoder input is the reliability vector z defined in 'Degraded-Factor Regimes'; that vector contains only K-derived statistics, a missingness indicator, a one-hot degradation tag, and the center-to-neighborhood-mean difference. It does not contain R, LS, or C. RASPL's formula branch injects R × K̃ × LS × C, and the formula-feature baseline receives log(1 + A_formula), so both have access to the product. Direct-MLP-STATS, by contrast, cannot see the three factors that dominate target variation, making its R2_all ≈ 0.10 in Table 2 an expected information floor, not a measure of direct prediction quality. A direct model given center R, LS, C, and degraded K̃ would likely score far higher. The abstract's claim that RASPL 'substantially outperforms matched direct prediction' therefore rests on an unfair baseline in the MLP comparison. The CNN comparison does provide all four channels, but the headline 0.733 R2 gain comes from the MLP table and is not a matched information comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies factor-derived proxy supervision, where the learning target is constructed from a known scientific equation rather than from independent observations. Using a RUSLE-derived soil-loss proxy, the authors introduce a diagnostic framework with degraded-formula references, tree-based baselines, matched direct and formula-feature predictors, contextual ablations, tail metrics, and a degradation robustness score (DRS). They then propose RASPL, a residual method that keeps the degraded formula estimate as a prediction anchor and learns an adaptively gated contextual correction. The central empirical claim is that formula preservation substantially outperforms matched direct prediction and provides a better degradation--tail tradeoff than treating the formula estimate as an ordinary input feature, thereby establishing formula preservation as a design principle. The paper includes controlled degradation regimes (noise, coarsening, masking, and complete K removal), three seeds, matched comparisons, and a separate missing-K stress test, and it explicitly discloses that no validation against independent soil-loss measurements is performed.","tokens_in":12238,"tokens_out":3523,"duration_ms":45849,"significance":"The paper addresses a real and underappreciated failure mode: high accuracy on factor-derived proxy targets may simply reconstruct the generating equation. The proposed diagnostic framework and the residual anchored prediction idea are potentially useful to the scientific-ML community. The authors should be credited for the controlled degradation protocol, the matched formula-feature and direct baselines, the separate treatment of the missing-K regime, the pre-specified visualization tile selection rule, and the explicit disclosure that validity against independent soil-loss measurements is not assessed. However, the central empirical comparison is undermined by an information-access asymmetry in the statistical-encoder baseline, and several claims rely on small differences across only three seeds without reported variance. Because the core contribution is the formula-preservation principle, these issues are load-bearing and require revision before the results can be accepted as stated.","major_comments":[{"comment":"The headline comparison of RASPL-MLP-STATS against Direct-MLP-STATS is not information-matched. In the section on matched baselines, direct models are said to retain the same contextual encoders and regime-matched inputs as their RASPL counterparts, but for MLP-STATS the contextual encoder input is the reliability vector z described in Degraded-Factor Regimes. That vector contains K-derived statistics, a missingness indicator, a one-hot degradation tag, and the center-to-neighborhood-mean difference; it does not contain R, LS, or C. RASPL's formula branch injects R times K-tilde times LS times C, and the formula-feature baseline receives log(1 + A_formula), so both have access to the product that dominates target variation. Direct-MLP-STATS cannot see R, LS, or C, which makes its R2_all close to 0.10 in Table 2 an expected information floor rather than a measure of direct prediction quality. The abstract's claim that RASPL 'substantially outperforms matched direct prediction' therefore rests on an unfair baseline in the MLP comparison. A direct MLP baseline that receives the center-pixel R, LS, C, and degraded K channels (or the same four-factor window) must be added and compared; until then the 0.733 R2 gain in Table 3 cannot be attributed to formula preservation.","section":"Matched Direct and Formula-Feature Baselines; Table 2; Table 3"},{"comment":"The key differences between RASPL and the formula-feature baseline are small: mean R2_all improvement of 0.015, Tail95 MAE improvement of 0.061, and only 8/9 cells improved in each metric, with Tail95 underprediction improved in only 5/9 cells. All results are macro-averaged over three seeds with no standard deviations, per-seed values, or significance tests reported. Given that the central design-principle claim hinges on these small advantages, the manuscript needs per-seed or interval estimates to show that the advantage is not within seed noise. This is particularly important because DRS is a normalized composite whose values depend on the comparison pool.","section":"Results and Analysis, Formula Preservation; Tables 2 and 3"},{"comment":"The degradation robustness score is a weighted, min-max normalized composite with weights 0.35, 0.25, 0.25, and 0.15, and the authors correctly note that DRS values from different normalization pools are not directly comparable. However, the paper's broader wording, such as 'stronger degradation robustness' and the ranking statements in the abstract, does not always carry this caveat. Since the component metrics are also reported, the ranking claims should be tied explicitly to those components, and ideally the sensitivity of the ranking to the chosen DRS weights should be examined, because the arbitrary weighting scheme could change the ordering of MLP-STATS versus CNN-RAW+STATS.","section":"Diagnostic Criteria, Eq. (6); Tables 2 and 4"}],"minor_comments":[{"comment":"Under K full the formula reference attains R2_all = 1.0000 by construction, as the authors acknowledge; this is a sensible reconstruction diagnostic, but the wording in Table 1's discussion could more clearly separate this identity check from model-based performance.","section":"Problem Setup and Diagnostics, Eq. (4)"},{"comment":"The manuscript reports results only for three seeds and does not state whether a single data split is used across seeds or whether the split itself is reseeded; reporting the split construction would clarify the matched comparisons.","section":"Experimental Setup, Evaluation Regimes"},{"comment":"The missing-K case shows that the Tail95 underprediction rate remains 0.9732 even for the best model, which is a strong caveat to the overall robustness claim; this caveat is disclosed, but it deserves a sentence in the abstract or conclusions so that the reader does not overgeneralize the robustness result.","section":"Results and Analysis, Missing-Factor Stress Test; Figure 2"},{"comment":"The paper would benefit from a data and code availability statement, since the reproducibility of the three-seed comparisons and the pre-specified visualization rule is otherwise hard to verify.","section":"Throughout"},{"comment":"Some reference entries have inconsistent formatting, such as 'V .M., P.' and spacing in author initials; a final copyediting pass is needed.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The core idea is worth publishing if the information-matched baseline issue is fixed and the statistical uncertainty is addressed. The current Table 2 comparison, as written, does not support the abstract's claim of 'substantially outperforms matched direct prediction' for the MLP case. The CNN comparison is more defensible but shows only a moderate R2 advantage over direct prediction, so the central design-principle claim needs to be rebalanced. The paper otherwise fits the journal's scope and the authors have been transparent about the absence of independent soil-loss validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: there's a real contribution here—the diagnostic protocol and the equation-reconstruction framing—but the paper's headline result is built on an unfair baseline. The direct MLP baseline is not information-matched: it never sees R, LS, or C, while RASPL gets them through the formula branch. That makes the 0.733 R2 gain over \"direct\" an access-to-information artifact, not a demonstration of formula preservation. The properly matched CNN comparison shows a real but smaller gain (0.34 R2), and the difference between RASPL and formula-as-feature is modest (0.015 R2 for MLP, and actually slightly negative R2 for CNN). So the central claim as stated overreaches.\n\nWhat is genuinely new and good: the paper formalizes \"equation reconstruction\" as a failure mode in factor-derived proxy supervision, and builds a sensible diagnostic suite—controlled degradation regimes, formula references, tree baselines, tail metrics, and a degradation robustness score. The RASPL design is a reasonable instantiation of residual learning around a known estimator, and the ablations (statistical vs CNN encoders, window sizes) are thorough and honestly reported. The authors also state clearly that they do not validate against independent soil-loss measurements, which is the right kind of honesty.\n\nThe soft spots are the ones you'd expect from a single-case-study paper with a somewhat circular setup. The tables report means without standard deviations or significance tests; three seeds is enough for a pilot, not for a strong claim. No code or data are released, which makes the \"matched\" claim hard to verify. The DRS is normalized within pools, so cross-table DRS values are not comparable (they do say this). The K-missing stress test is reported separately and properly caveated. All of these are addressable.\n\nThe bigger issue is the information-matching. The authors call the baselines \"matched\" but they are matched in encoder architecture, not in input access. This needs to be fixed before the paper's central comparison is credible. The CNN direct baseline with all four channels is the right comparison, and the gain there is much smaller than the headline. The paper should be revised to lead with that.\n\nWho this is for: researchers building ML systems on factor-derived proxy targets (environmental ML, remote sensing, physics-informed learning) will find the diagnostic protocol useful. The specific RASPL architecture is less important than the evaluation discipline.\n\nRecommendation: engage with it. Send it to serious peer review, but expect major revision. The novelty is real, the execution is mostly careful, and the flaws are fixable. If the authors fix the baseline matching, add uncertainty estimates, and release code/data, this could be a solid contribution.","headline":"Useful diagnostic protocol, but the headline result rests on an information-mismatched baseline; fix that and this deserves a serious referee.","tokens_in":12762,"tokens_out":3393,"would_cite":true,"duration_ms":39290,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"When proxy targets come from known equations, keeping the formula as an anchor beats direct prediction.","keywords":["proxy supervision","RUSLE","equation reconstruction","residual learning","degradation robustness","soil erosion","scientific machine learning","RASPL"],"falsifier":"Compare RASPL against direct prediction using independently measured field soil-loss values under degraded K inputs; if direct prediction matches or exceeds RASPL's accuracy on observed measurements, the claim that formula preservation is the central design principle for robust proxy learning is falsified.","tokens_in":11756,"feed_emoji":"🧮","tokens_out":4895,"duration_ms":53175,"temperature":0.7,"pith_summary":"The paper studies what happens when machine-learning targets are computed from known scientific formulas rather than measured. It shows that when the same factors are used as inputs, high accuracy can simply mean the model re-learned the formula, not that it is robust to degraded factor information. The authors introduce diagnostics to separate these cases and propose RASPL, which keeps the degraded formula estimate as a fixed anchor and learns a gated contextual correction. In RUSLE-based soil-loss prediction, RASPL outperforms direct prediction and beats treating the formula as an ordinary input feature on robustness and tail error. The paper concludes that formula preservation is the central design principle for learning from factor-derived proxy targets.","feed_headline":"Formula-anchored residuals beat direct prediction on proxy targets","feed_subtitle":"When soil-loss targets come from RUSLE, keeping the degraded formula as a fixed anchor improves robustness and tail error.","key_machinery":"RASPL combines a fixed formula anchor and an adaptive gate in log space: the predicted log value is clip(y_formula + $\\alpha$ * delta, -5, 5), where delta is a contextual residual proposal and $\\alpha$ = 1 - $\\sigma$(logit) is a learned gate value conditioned on a reliability vector that encodes factor degradation. A residual penalty, q * $delta^{2}$, discourages large corrections when the gate favors retaining the formula. The formula anchor is the degraded RUSLE estimate A = R * K_tilde * LS * C, kept as a reference rather than treated as an ordinary input feature.","core_discovery":"The central claim is that in factor-derived proxy supervision, a model that predicts the proxy directly from the same factors can appear excellent while merely reconstructing the generating equation, which is not robustness. The paper's RASPL framework addresses this by defining the prediction as the degraded formula estimate plus an adaptively gated residual learned from context, in log space. Matched experiments on a RUSLE soil-loss proxy show RASPL-MLP-STATS reaches $R^{2}$_all 0.8343 versus 0.1017 for direct MLP, with the formula-feature baseline at 0.8188; the CNN variant gives the best Tail95 MAE and degradation robustness. The paper states that these results establish formula preservation as the central design principle for robust learning from factor-derived proxy targets.","pith_inferences":["The same reconstruction-versus-robustness ambiguity likely applies to other factor-derived proxies such as evapotranspiration formulas, carbon-cycle models, or climate indices, so the diagnostic and residual-anchor design could transfer.","Because the paper does not validate against independently observed soil-loss measurements, a natural testable extension is to run the matched RASPL-versus-direct comparison against field-measured erosion data.","The K-missing stress test shows that even with strong contextual correction, complete factor absence leaves a high underprediction rate; imputing K from neighboring factors instead of a constant mean fallback could be a follow-up."],"forward_implications":["When the supervision target is generated by a known equation with possibly degraded inputs, models should anchor on the formula estimate rather than predict the target from scratch.","Adding the formula as a standard input feature captures most of the accuracy gain, but explicit preservation further improves tail robustness and degradation robustness.","Neighborhood statistics suffice for average accuracy, while convolutional context improves tail robustness, suggesting different encoder choices for different error regimes.","Larger convolutional windows do not help; a 3x3 window is sufficient, so additional computational cost is not justified by accuracy gains."],"supporting_citations":[{"why":"Reviews the (Revised) Universal Soil Loss Equation, establishing the proxy target's empirical basis.","marker":"Benavidez et al. 2018"},{"why":"Provides the European soil-loss assessment that motivates RUSLE-derived proxy targets.","marker":"Panagos et al. 2015"},{"why":"Assesses global soil erosion with RUSLE-derived estimates, grounding the problem's importance.","marker":"Borrelli et al. 2017"},{"why":"Uses RUSLE with remote sensing and GIS in an arid zone, a representative proxy-target application.","marker":"Li et al. 2023"},{"why":"Introduces AI geospatial layers into RUSLE, the type of ML-enhanced formula use the paper compares against.","marker":"Samarinas et al. 2024"},{"why":"Combines the RUSLE framework with machine learning, the case-study setting the paper builds on.","marker":"Ge et al. 2023"},{"why":"Supplies the corruption-robustness evaluation lineage that the degradation regimes extend.","marker":"Hendrycks and Dietterich 2019"},{"why":"Physics-informed neural networks, the approach contrasted with formula-preserving prediction.","marker":"Raissi, Perdikaris, and Karniadakis 2019"}],"fun_headline_variants":["Formula anchor beats direct prediction for proxy targets","Proxy prediction? Keep the formula, fix the residuals","Why your proxy model just reconstructs the formula","RASPL: improve proxy targets without ditching the formula","Direct proxy predictions fail; formula-anchored residuals win"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes the target raster is exactly the RUSLE product of the four factors and that degrading only K is the right test of model quality; the paper explicitly does not assess validity against independently observed soil-loss measurements.","fun_headline_variants_meta":{"raw":{"variants":["Formula anchor beats direct prediction for proxy targets","Proxy prediction? Keep the formula, fix the residuals","Why your proxy model just reconstructs the formula","RASPL: improve proxy targets without ditching the formula","Direct proxy predictions fail; formula-anchored residuals win"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000193,"raw_usage":{"total_tokens":1337,"prompt_tokens":916,"completion_tokens":421,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":345}},"tokens_in":532,"tokens_out":421,"duration_ms":5960,"temperature":1.0,"reasoning_tokens":345,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:48:26.031993+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare RASPL against direct prediction using independently measured field soil-loss values under degraded K inputs; if direct prediction matches or exceeds RASPL's accuracy on observed measurements, the claim that formula preservation is the central design principle for robust proxy learning is falsified.","supporting_citations":[{"cited_title":"and Jackson, B","cited_arxiv_id":null,"evidence_quote":"Reviews the (Revised) Universal Soil Loss Equation, establishing the proxy target's empirical basis."},{"cited_title":"and Kalopesa, Eleni and Zalidis, George C","cited_arxiv_id":null,"evidence_quote":"Introduces AI geospatial layers into RUSLE, the type of ML-enhanced formula use the paper compares against."}],"review_version":1}