{"id":"db5747fd-948a-460f-aaa7-a66ae210779b","arxiv_id":"2605.28076","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.5,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Squared-loss scientific surrogates hit a conditional-mean barrier that erases task-relevant variability; residual-feature orthogonality and effect-size diagnostics detect it, and richer likelihood losses are required to recover it.","lead":"Squared-loss surrogates in science often learn only the average answer and erase the variability that coarse-grained systems still need. The paper names this conditional-mean barrier and gives diagnostics plus a loss prescription so modelers can tell when a point predictor is the wrong object.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged diagnostic-generalization limit.","rationale":"The strongest claim rests on a standard fact of L2 regression (MSE recovers E[Y|X]) plus a transparent consequence for stochastic models under the same paired loss, both of which hold under the paper's stated assumptions. The diagnostic framework and the modeling prescription follow directly. The numerical illustrations are consistent with that logic. The reader's weakest assumption correctly isolates the remaining practical risk: that the chosen residual features and effect-size cut-offs actually track scientifically relevant variability rather than merely residual scatter relative to those features. Because that risk is already reflected in the CONDITIONAL verdict and moderate confidence, no further downward adjustment is required. A single independent-feature recomputation on the Lorenz-96 example would settle whether the diagnostic transfers or remains example-specific.","tokens_in":12206,"tokens_out":464,"duration_ms":4511,"concrete_test":"On the Lorenz-96 closure, recompute residual-feature orthogonality and effect sizes after replacing the paper's residual features with an independent, task-derived set (e.g., energy-spectrum bins or multi-step rollout variance of the slow variables). If the barrier diagnosis flips or the effect-size ranking changes materially, the operational reliability of the diagnostic is weaker than claimed; if it is stable, the reader's concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is classical and internally secure: under paired squared loss the L2 risk minimizer is the conditional mean, so deterministic (and even stochastic) predictors trained that way cannot recover non-mean features of the conditional law; residual-feature orthogonality plus effect-size then separates that irreducible scatter from deterministic misspecification. The two numerical studies (two-branch law and two-scale Lorenz-96 closure) illustrate the mechanism and the loss prescription. The only soft spot is precisely the one the reader already named: whether the particular residual features and effect-size thresholds used on those two examples constitute a reliable operational test for task-relevant variability on broader SciML surrogates. That is a scope/transfer limitation, not an internal inconsistency or a flaw in the math. No stronger load-bearing attack is warranted.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper formulates the conditional-mean barrier: after coarse-graining or partial observation, many SciML prediction tasks are one-to-many, so deterministic surrogates trained by paired squared loss recover the conditional mean E[Y|X] while missing task-relevant variability in the conditional law. It proposes a diagnostic that combines residual-feature orthogonality with effect-size measures to separate deterministic underfitting from irreducible conditional scatter, and notes that stochastic predictors under the same paired MSE objective are driven back to the conditional mean because model variance is penalized. The modeling prescription is that when residual variability matters, the loss must score richer features of the conditional law (e.g., likelihood-based stochastic-scale models). Reproducible studies on a controlled two-branch law and a two-scale Lorenz-96 closure illustrate barrier identification, suppression of collective fluctuation statistics under deterministic closures in rollout, and partial recovery of variability by a minimal stochastic-scale model.","tokens_in":12327,"tokens_out":1199,"duration_ms":18628,"significance":"If the diagnostic is used as intended, the paper gives SciML practitioners a concrete way to decide when a well-fitting deterministic surrogate is still the wrong object for the scientific task, and a clear loss-level prescription rather than an architecture-level one. The core L2 fact (population minimizer is the conditional mean; paired MSE penalizes predictor variance) is classical, but naming the barrier, packaging residual-feature orthogonality with effect-size, and demonstrating the rollout consequence on a two-scale Lorenz-96 closure are useful contributions for the community. Strengths include an explicit, non-circular definition of the barrier from the known population minimizer, a falsifiable residual diagnostic, and reproducible numerical studies. The main limitation is transfer: whether the chosen residual features and effect-size thresholds reliably flag task-relevant variability beyond the two examples.","major_comments":[{"comment":"The diagnostic framework (residual-feature orthogonality plus effect-size) is load-bearing for the claim that one can operationally identify the barrier rather than only restate the L2 fact. The residual feature bank and effect-size thresholds are free parameters of the method. The manuscript needs a clearer statement of how those features are chosen for a new problem, what constitutes a decisive effect size, and how the procedure avoids mistaking structured misspecification for irreducible conditional variability. Without that guidance or a sensitivity study, the two numerical examples remain illustrations rather than a transferable operational test.","section":"Diagnostic framework / residual-feature orthogonality"},{"comment":"The phrase task-relevant variability is central to the modeling prescription but is not given an independent operational definition beyond residual scatter relative to the chosen features. In the Lorenz-96 closure study, the paper should state explicitly which collective fluctuation statistics are the scientific target (e.g., variance, spectra, extremes of the resolved variables in rollout) and show that the residual features used in the diagnostic are aligned with those targets, not only with one-step residual magnitude. Otherwise the claim that the diagnostic identifies scientifically important missing variability rests on the same free feature choice.","section":"Two-scale Lorenz-96 closure / rollout statistics"},{"comment":"The consequence that stochastic outputs under paired squared loss do not overcome the barrier is correct but standard (the excess risk includes the model variance term). The manuscript should more carefully separate this pedagogical clarification from the novel contribution (the residual diagnostic and the SciML prescription). As written, the abstract and introduction can be read as presenting the stochastic-under-MSE fact as a primary result; tightening the novelty claim would strengthen the paper without changing the math.","section":"Abstract / introduction; paired squared-loss argument"}],"minor_comments":[{"comment":"Define residual-feature orthogonality and the effect-size statistic with explicit formulas early (population and sample versions), and keep notation consistent between the two-branch and Lorenz-96 sections.","section":"Diagnostic framework"},{"comment":"State training protocol, architecture, and hyperparameter choices for the deterministic and stochastic surrogates in enough detail for independent reimplementation; point to the reproducibility package if that is the intended source of truth.","section":"Numerical studies"},{"comment":"In the two-branch law example, report both residual diagnostics and a simple visualization of the recovered conditional law (or lack thereof) so readers can see the barrier without relying only on scalar effect sizes.","section":"Controlled two-branch law"},{"comment":"Clarify whether the likelihood-based stochastic-scale model is trained with a proper scoring rule for the full conditional (or a parametric scale family) and how that differs from injecting noise into an MSE-trained mean model.","section":"Stochastic-scale model"},{"comment":"Minor polish: ensure figure captions state what is being compared (truth vs deterministic closure vs stochastic-scale) and what time horizon the rollout statistics cover.","section":"Figures / rollout panels"}],"recommendation":"minor_revision","confidential_remarks":"The mathematical core is standard and correctly applied; the value is in framing and diagnostics for SciML. I do not see a load-bearing error that would justify reject. The main risk is overclaiming the diagnostic as a general operational test on the strength of two examples with free feature/threshold choices. If the authors qualify scope and add choice guidance or sensitivity, minor revision is appropriate. Fit for a stat.ML / SciML methods venue is good; less so for a pure theory journal."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline: this paper does not invent new probability. It names a real practice failure in scientific surrogates—squared-loss models learn E[Y|X] and can miss task-relevant conditional variability—and gives a residual diagnostic plus a loss prescription that people in closures actually need.\n\nWhat is new is the packaging and the operational workflow. Residual-feature orthogonality plus effect-size is a clean way to separate deterministic underfit from irreducible scatter. The paper is also explicit that stochastic heads under paired MSE do not fix the problem, because the objective penalizes model variance and drives you back to the mean. That consequence is standard but often ignored in SciML; stating it and then showing a minimal likelihood-style alternative on a two-branch law and a two-scale Lorenz-96 closure is useful. The math core is correct, circularity is low, and the modeling prescription (score richer features of the conditional law when variability matters) is the right one.\n\nSoft spots are real but proportional. Novelty is moderate: the L2 fact is classical. The diagnostic’s feature bank and effect-size thresholds are only demonstrated on two controlled examples, so transfer to broader surrogates is the open question—not an internal contradiction. Empirical breadth is thin; reproducibility is claimed rather than verified from the extract. None of that sinks the central claim.\n\nThis is for multi-scale modeling and SciML people who still treat low MSE as success when fluctuation statistics matter. A serious editor should send it to referees. I would cite the barrier framing when discussing MSE-trained closures, and I would put it on a reading-group list if the room cares about surrogate losses. Engage with the work; it is solid methodological SciML, not a breakthrough and not fluff.","headline":"Classical MSE→conditional-mean fact packaged as a usable SciML diagnostic and loss prescription; moderate novelty, clear argument, thin but honest empirics.","tokens_in":12997,"tokens_out":450,"would_cite":true,"duration_ms":12065,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Squared-loss surrogates hit a conditional-mean barrier that erases the variability science needs.","keywords":["conditional-mean barrier","scientific machine learning","surrogate models","squared loss","stochastic closures","Lorenz-96","residual diagnostics","conditional law"],"falsifier":"On a new one-to-many scientific surrogate, compute residual-feature orthogonality and effect size; if residuals remain large and nearly orthogonal yet a richer likelihood or proper scoring rule still fails to recover the missing fluctuation statistics that the paper’s Lorenz-96 experiment recovered, the diagnostic-plus-prescription claim fails.","tokens_in":13005,"feed_emoji":"📉","tokens_out":676,"duration_ms":5645,"temperature":0.7,"pith_summary":"In many scientific prediction tasks, coarse graining or partial observation turns a one-to-one map into a one-to-many conditional law: the same input can produce a cloud of plausible outputs. Deterministic machine-learning surrogates trained by ordinary squared loss still converge to a mathematically well-defined object—the conditional mean—but that single number discards the scatter that often carries the physically relevant fluctuations. The paper names this the conditional-mean barrier and supplies a practical diagnostic: residual-feature orthogonality checks whether the leftover error is uncorrelated with the features that matter, while effect-size measures decide whether that leftover scatter is large enough to matter for the task. The same analysis shows why simply making the network stochastic does not fix the problem under paired squared loss—the objective itself penalizes model variance and drives every realization back to the mean. When the diagnostic flags the barrier, the prescription is clear: switch to a loss that scores richer features of the conditional distribution. Controlled two-branch examples and a two-scale Lorenz-96 closure demonstrate that the diagnostic correctly identifies the barrier, that deterministic closures suppress collective fluctuation statistics in long rollouts, and that a minimal likelihood-based stochastic model can restore much of the missing variability.","feed_headline":"Squared-loss surrogates erase the variability science needs","feed_subtitle":"A residual diagnostic flags the conditional-mean barrier and shows why stochastic nets alone do not fix it.","key_machinery":"The residual-feature orthogonality and effect-size diagnostic: after a surrogate is fit, residuals are tested for orthogonality to chosen residual features; large residual effect size with near-orthogonality signals irreducible conditional variability rather than under-fitting, diagnosing the barrier.","core_discovery":"Deterministic surrogates trained by paired squared loss learn the conditional mean and therefore miss task-relevant variability in the underlying conditional law—the conditional-mean barrier. Residual-feature orthogonality together with effect-size diagnostics can separate this irreducible conditional scatter from ordinary deterministic underfitting, and the same squared-loss objective forces even stochastic predictors back to the mean.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Squared loss locks SciML surrogates to the conditional mean","Residual diagnostics flag the conditional-mean barrier","Even stochastic nets hit the mean under paired squared loss","How to separate underfitting from irreducible conditional scatter","Conditional-mean barrier erases task-relevant variability"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That the residual features and effect-size thresholds chosen on the two numerical examples actually capture the variability that matters for the scientific task, rather than merely residual scatter relative to those particular features.","fun_headline_variants_meta":{"raw":{"variants":["Squared loss locks SciML surrogates to the conditional mean","Residual diagnostics flag the conditional-mean barrier","Even stochastic nets hit the mean under paired squared loss","How to separate underfitting from irreducible conditional scatter","Conditional-mean barrier erases task-relevant variability"]},"model":"grok-4.5","effort":"low","cost_usd":0.00533,"raw_usage":{"total_tokens":1459,"prompt_tokens":760,"num_sources_used":0,"completion_tokens":57,"cost_in_usd_ticks":53300000,"prompt_tokens_details":{"text_tokens":760,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":642,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":760,"tokens_out":57,"duration_ms":4880,"temperature":1.0,"reasoning_tokens":642,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T15:49:48.053587+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a new one-to-many scientific surrogate, compute residual-feature orthogonality and effect size; if residuals remain large and nearly orthogonal yet a richer likelihood or proper scoring rule still fails to recover the missing fluctuation statistics that the paper’s Lorenz-96 experiment recovered, the diagnostic-plus-prescription claim fails.","supporting_citations":[],"review_version":2}