{"id":"23514e42-0645-48dc-b0e9-24966b275a89","arxiv_id":"2412.06110","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Multifidelity estimators can quantify Greenland ice mass loss uncertainty at a fraction of Monte Carlo cost, but the headline speedups depend on free access to a covariance matrix from pilot runs.","lead":"This paper applies three existing multifidelity uncertainty quantification methods to Greenland ice sheet simulations, estimating the expected 2015-2050 ice mass loss under uncertain basal friction and geothermal heat flux. It reports that the best method reaches a target accuracy about 91 times faster than plain Monte Carlo, though the comparison excludes the cost of estimating the model correlations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pilot-cost accounting: the reported 12.6x/10.2x/91.0x speedups exclude the ~199 CPU-days used to estimate the 13x13 covariance matrix (Sec. 4.1); charging that overhead once reduces the speedups to roughly 2x, so the abstract's 'two orders of magnitude' claim is not yet supported.","rationale":"The central claim is a computational speedup claim, so the denominator and numerator of the speedup must be apples-to-apples. The paper's own numbers show the excluded pilot stage costs roughly half of the Monte Carlo comparison budget; this is not a rounding error. The reader's weakest assumption identifies exactly this issue, and my estimate confirms it: including pilot costs collapses the speedups from 12.6/10.2/91.0 to roughly 1.7/1.6/1.9, or at most about 1.9 even if every optimized sample is reused from the pilot. I therefore agree with the reader, and I would keep the CONDITIONAL verdict rather than moving to ACCEPT or REJECT. The paper has real strengths: it uses a community code, follows ISMIP6 protocol, provides the code, and verifies estimators against Monte Carlo. The concern is not about the unbiasedness or mathematical correctness of MFMC/MLMC/MLBLUE, nor about the authors' integrity; it is about the economic comparison underlying the headline. The body does disclose the exclusion, but the abstract's unqualified 'two orders of magnitude' overstates what is demonstrated. A conditional acceptance with a required revision of the cost accounting and headline is the appropriate outcome.","tokens_in":26248,"tokens_out":12800,"duration_ms":125650,"concrete_test":"Using the released repository, recompute the target-accuracy speedups with total cost C_total = C_pilot + C_est - C_overlap, where C_pilot is the cost of all runs used to construct Sigma and the surrogates (32 evaluations of s1-s6, 128 of s7-s12, 50,000 of s13, and 9 SSA-coarse reference runs, priced at Table 2), C_est is the optimized budget reported in Section 4.2, and C_overlap is the cost of optimized samples already contained in the pilot set. Recompute the speedups relative to the 381.22 CPU-day Monte Carlo cost. If, as the arithmetic above indicates, all speedups fall below 2x, revise the abstract and Section 4.2 to state explicitly that speedups exclude the covariance-estimation and surrogate-construction stage, and reframe the headline claim as conditional on the availability of free pilot data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 and Section 4.1 explicitly exclude the cost of computing the 13x13 covariance matrix Sigma; Section 4.1 states 'we do not count these computations towards our budget.' This exclusion is load-bearing because the pilot stage is large: 32 HO pilot samples (s1-s6) cost roughly 194 CPU-days, and 128 SSA samples (s7-s12) add roughly 5 CPU-days, before the 50,000 s13 evaluations and 9 reference SSA-coarse runs used to build the interpolation/extrapolation surrogates. Thus the pilot stage alone costs about 199 CPU-days, roughly half the 381.2 CPU-day Monte Carlo target budget and 47x the reported 4.2 CPU-day MLBLUE budget. A user starting from scratch must spend at least this amount, and if pilot samples are reused in the estimator there may be little or no additional cost, so the speedup relative to MC is at most 381.2/199 ~ 1.9x, not 91.0x; MFMC and MLMC are similarly bounded. This is not a bookkeeping nit: the headline 'two orders of magnitude' in the abstract is unsupported by the fair from-scratch comparison, and the body's largest reported factor is 91.0, not 100.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper applies three multifidelity Monte Carlo estimators—MFMC, MLMC, and MLBLUE—to quantify uncertainty in projected 2015–2050 Greenland ice mass loss under parametric uncertainty in basal friction and geothermal heat flux. The high-fidelity model is an ISSM higher-order ice-flow model; lower-fidelity surrogates are obtained by mesh coarsening, SSA physics reduction, time-truncation/extrapolation, and interpolation. The estimators are standard and are presented with algorithms and MSE formulas. The numerical demonstration reports that, for a target accuracy of ±1 mm SLR at 95% confidence, the multifidelity estimators require 30.1, 37.5, and 4.2 CPU-days versus 381.2 CPU-days for Monte Carlo, corresponding to speedups of 12.6, 10.2, and 91.0, which the abstract summarizes as two orders of magnitude. However, the 13×13 model covariance matrix is estimated from 32 HO pilot samples plus 96 additional SSA samples, and Section 4.1 explicitly excludes this pilot cost from all budgets; this exclusion is the main driver of the headline speedups. The paper also treats the estimated covariance as known in the MSE and confidence-interval calculations.","tokens_in":26551,"tokens_out":10953,"duration_ms":107568,"significance":"If the numerical claims held under a fair from-scratch cost accounting, this would be a valuable demonstration that unbiased multifidelity UQ is tractable for continental-scale ice sheet simulations. The paper's strengths are its use of a community code (ISSM), its adherence to the ISMIP6 protocol, its comparison against Monte Carlo estimates, and its provision of code and prediction data. The theoretical estimators are standard and correctly presented, and the unbiasedness guarantees are not in question. However, the headline quantitative claims currently depend on excluding a pilot cost that is comparable to the Monte Carlo target budget, so the demonstrated significance is not yet established. The methodological contribution is incremental rather than radically new, but the application is novel and potentially useful to the ice-sheet modeling community.","major_comments":[{"comment":"Section 3.1 states that estimation of the model covariance matrix Σ is not counted toward computational costs, and Section 4.1 repeats this for the pilot computations ('we do not count these computations towards our budget'). This exclusion is load-bearing for the central speedup claim. Using the costs in Table 2, the 32 HO pilot runs (models s1–s6) cost about 193.7 CPU-days, and the additional 96 SSA runs (models s7–s12) cost about 3.9 CPU-days, for a pilot stage of roughly 199 CPU-days—about half of the 381.2 CPU-day Monte Carlo budget and about 47 times the reported 4.2 CPU-day MLBLUE budget. A from-scratch user must spend at least this amount, and even under the paper's reuse suggestion the total cost is at least about 199 CPU-days, bounding any speedup by roughly 381.2/199 ≈ 1.9. The reported speedups of 12.6, 10.2, and 91.0, and the abstract's 'two orders of magnitude,' are therefore not supported by the demonstration as accounted. The paper should either include pilot cost in a from-scratch comparison or explicitly rescope the claims to 'incremental cost given a pre-existing covariance estimate.'","section":"Sections 3.1 and 4.1"},{"comment":"The covariance matrix Σ is estimated from only 32 samples for all entries involving HO models, yet the MSE formulas (11), (17), and (26), together with the optimal sample-size rules (14)–(15) and the MLBLUE SDP (27), treat this estimate as the true covariance. With n = 32, the sampling error in a correlation estimate is on the order of 0.15–0.2, and the MFMC allocation depends on differences of squared correlations while the MLBLUE optimization is known to be sensitive to approximation errors in Σ (as the authors themselves note in Section 4.4). The paper does not report a sensitivity or bootstrap analysis, nor does it verify in repeated experiments that the MSE-predicted target accuracy is achieved. Since the 'guaranteed target accuracy' is the basis for the reported speedups, this finite-sample uncertainty should be quantified.","section":"Sections 3.1, 4.1, and 4.2"},{"comment":"The 95% confidence intervals in Table 3 are computed from the estimator variance together with normal quantiles. For the target-accuracy runs, the high-fidelity model is sampled 7 times by MFMC, 12 times by MLMC, and only once by MLBLUE (Section 4.2 and Figure 8). The ice mass loss distribution is explicitly shown to be non-Gaussian in Section 2.2 and Figure 4, so the normal approximation for the MLBLUE estimator with a single high-fidelity sample is not justified by the CLT, and the claimed 95% confidence level may not hold. The paper should either use conservative inequalities, provide a bootstrap or coverage study, or restrict the confidence statements to settings with larger high-fidelity sample sizes.","section":"Sections 4.2 and 4.3, Table 3"}],"minor_comments":[{"comment":"The abstract's claim of 'two orders of magnitude' speedup overstates the largest factor reported in the body, 91.0×; the body itself says 'between one and two orders of magnitude.' Please align the abstract with the quantitative results.","section":"Abstract"},{"comment":"The ocean water density is given as 'ρw = 1023 km/m3'; the unit should be kg/m3.","section":"Section 2.1.2"},{"comment":"The citation to Paterson appears as '[ ?, p. 97]]Paterson1994' and needs to be corrected or completed.","section":"Section 2.1.2"},{"comment":"The symbol Σ is used both for the model covariance matrix and for the diagonal covariance matrix Σα of the basal friction coefficients; consider renaming one to avoid confusion.","section":"Section 3.1"},{"comment":"There is a duplicated word in 'robust against against approximation bias'; it should read 'robust against approximation bias.'","section":"Section 4.4"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a genuinely useful application paper—the first head-to-head of MFMC, MLMC, and MLBLUE on a realistic Greenland ISSM setup with purpose-built surrogates—but the abstract's 'two orders of magnitude' speedup is not supported as stated. The body reports 12.6x, 10.2x, and 91.0x, and those numbers exclude the ~199 CPU-days spent estimating the 13x13 covariance matrix from 32 HO and 128 SSA pilot runs (Sec. 4.1). Charging that once drops the from-scratch speedups to roughly 1.6–1.9x. That's a real load-bearing flaw, not a bookkeeping nit.\n\nWhat's good: The method exposition is clear and the algorithms are complete. The surrogate construction is thoughtful—mesh coarsening, SSA physics, extrapolation, and an ultra-cheap interpolation model—and the paper shows quantitatively that naive surrogate use is badly biased (Fig. 5) while the multifidelity estimates are unbiased and consistent with Monte Carlo. Code and data are public. The comparison of the three estimators on a real problem is new and informative: MLBLUE does win, as theory predicts, but MFMC and MLMC are more robust to covariance error, which the authors note. I believe the central methodological claim—that multifidelity UQ is practical for continental-scale ice sheets—holds up once you treat the pilot cost as part of the overhead.\n\nThe soft spots: (1) The pilot-cost exclusion is transparently disclosed but not adequately flagged in the abstract; a reader would reasonably take 'two orders of magnitude' at face value. (2) The covariance is estimated from 32 HO samples for a 13x13 matrix; that's noisy, and the paper doesn't quantify how estimation error in Sigma propagates into the reported MSE and speedups. They mention MLBLUE's sensitivity but don't test it. (3) The 91x MLBLUE number relies on a single high-fidelity sample and a near-perfect surrogate correlation; it's legal but brittle, and a referee should ask for sensitivity analysis on Sigma.\n\nVerdict: worth a serious referee. I'd send it out, but require the authors to either include pilot costs in a from-scratch comparison, or clearly amortize them over a stated number of UQ tasks. The paper is a good fit for a computational geoscience or UQ journal. I wouldn't cite the speedup numbers as they stand, but I'd point people to the method comparison and the surrogate setup.","headline":"Useful first head-to-head of multifidelity UQ estimators on a real Greenland model, but the headline 'two orders of magnitude' speedup excludes ~199 CPU-days of pilot covariance estimation and is not supported as stated.","tokens_in":27071,"tokens_out":5337,"would_cite":false,"duration_ms":50304,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["65C05","86A40"],"pacs":[],"model":"deepseek-v4-flash","headline":"Multifidelity estimators cut the cost of unbiased ice-sheet uncertainty quantification by up to 91-fold on a Greenland test case.","keywords":["multifidelity uncertainty quantification","ice sheet simulation","Greenland ice sheet","sea level rise","Monte Carlo estimation","best linear unbiased estimator","basal friction uncertainty","geothermal heat flux"],"falsifier":"Recount all speedups with the 32 high-fidelity and 96 low-fidelity pilot runs included in each estimator's budget, and compare the resulting cost with Monte Carlo at the same 1 mm sea-level target; if the multifidelity estimates no longer cost less, the headline speedup claim fails for that accounting.","tokens_in":26024,"feed_emoji":"🧊","tokens_out":5931,"duration_ms":56272,"temperature":0.7,"pith_summary":"This paper argues that formal uncertainty quantification for continental-scale ice-sheet simulations is achievable if the estimator is built from several models of different fidelity instead of from Monte Carlo sampling of the expensive high-fidelity model alone. It demonstrates the argument on a Greenland ice-sheet model whose basal friction and geothermal heat flux are treated as uncertain, estimating the expected 2015-2050 ice mass loss. The three multifidelity estimators---multifidelity Monte Carlo, multilevel Monte Carlo, and the multilevel best linear unbiased estimator---are all unbiased, so they target the high-fidelity answer rather than a biased surrogate answer. For a target accuracy of $\\pm1$ mm sea-level equivalent at 95% confidence, the paper reports speedups of $12.6\\times$, $10.2\\times$, and $91.0\\times$ over Monte Carlo, corresponding to about 30, 38, and 4 CPU-days instead of 381 CPU-days. The point is that unbiased uncertainty quantification, previously considered prohibitive at this scale, becomes computationally accessible.","feed_headline":"Multifidelity UQ cuts ice-sheet uncertainty cost 91-fold","feed_subtitle":"Unbiased estimators hit 1 mm sea-level accuracy in days instead of a year of supercomputer time.","key_machinery":"The machinery is the cross-model covariance matrix $\\Sigma$, the table of variances and correlations among the 13 model outputs, which all three estimators consume to allocate samples. Multifidelity Monte Carlo uses it to set regression weights $\\alpha_j = \\Sigma_{1j}/\\Sigma_{jj}$ and sample counts; multilevel Monte Carlo uses it to form level variances $\\mathrm{Var}(s_j - s_{j+1})$ and balance a telescoping sum; the multilevel best linear unbiased estimator casts estimation as a Gauss-Markov regression over every nonempty subset of models and solves a semidefinite program for sample sizes. The paper's particular surrogate ladder---higher-order and shallow-shelf physics on fine, medium, and coarse meshes, plus extrapolation and interpolation surrogates---keeps correlations with the high-fidelity model high enough and costs low enough that the allocation shifts nearly all sampling work onto cheap models.","core_discovery":"The paper's central claim is that the expected ice mass loss of a high-fidelity ice-sheet model can be estimated to a prescribed accuracy without paying the full Monte Carlo price, by exploiting correlation with cheaper models while preserving unbiasedness. On the Greenland test case, the best linear unbiased estimator uses one high-fidelity run, many surrogate runs of several types (coarser meshes, shallow-shelf physics, extrapolated and interpolated surrogates), and a covariance matrix estimated from pilot samples, to reach the 1 mm sea-level accuracy target in 4.2 CPU-days, whereas Monte Carlo needs 381.2 CPU-days. The estimates are verified against Monte Carlo through overlapping confidence intervals, and in a fixed low budget of five high-fidelity solves the best linear unbiased estimator produces a confidence interval as tight as the one Monte Carlo would need roughly 491 high-fidelity solves to match.","pith_inferences":["Beyond the paper: if the pilot runs used to estimate the covariance matrix (32 high-fidelity and 96 low-fidelity evaluations) were charged to the estimator budgets, the reported speedups would shrink by roughly an order of magnitude, and the best linear unbiased estimator's advantage over Monte Carlo for the target accuracy could become marginal.","Beyond the paper: the same covariance-allocation machinery should transfer directly to Antarctic projections and to other high-cost geophysical simulators, since the only new ingredients are a correlated surrogate family and a pilot estimate of the cross-model covariance matrix.","Beyond the paper: because the framework is agnostic to how uncertainties are characterized, the basal-friction and heat-flux uncertainty model could be replaced by posterior distributions from Bayesian inversions, converting the output from a prior-driven spread into a decision-relevant posterior predictive.","Beyond the paper: a cheap stress test of the headline claim would be to recompute the optimal sample allocations with a covariance matrix estimated from substantially more pilot samples and check whether the 91-fold speedup persists."],"forward_implications":["Unbiased uncertainty quantification for continental ice sheets becomes computationally accessible at the 1 mm sea-level scale: tens of CPU-days rather than roughly a CPU-year of Monte Carlo sampling.","The best linear unbiased estimator is guaranteed by construction to have variance no larger than the other two multifidelity estimators, and in the Greenland test it is the only method whose confidence interval at a five-high-fidelity-solve budget is comparable to hundreds of Monte Carlo runs.","Cheap surrogates that are poor approximations individually, such as the linear interpolation model, still pay off when they are highly correlated with the high-fidelity output, because the estimator shifts samples to them without introducing bias.","For a fixed small budget, all three multifidelity estimators produce tighter confidence intervals than Monte Carlo, with the best linear unbiased estimator reducing the interval width by more than a factor of nine at the tested budget.","The framework is agnostic to how the parametric uncertainties are characterized, so the same estimator structure applies to other ice-sheet codes, other regions, and other uncertain inputs without changing the algorithm."],"supporting_citations":[{"why":"Supplies the multifidelity Monte Carlo estimator with optimal weights and sample sizes that Section 3.2 implements.","marker":"[76]"},{"why":"Provides the multilevel Monte Carlo telescoping-sum formulation and sample-size balancing used in Section 3.3.","marker":"[28]"},{"why":"Introduces the multilevel best linear unbiased estimator whose variance the paper compares and extends.","marker":"[86]"},{"why":"Supplies the semidefinite-programming sample-size optimization that makes the best linear unbiased estimator practical for the 13-model set.","marker":"[19]"},{"why":"Is the community ice-sheet code used to run the high-fidelity and surrogate simulations.","marker":"[55]"},{"why":"Defines the projection protocol and 2015-2050 setup that the experiments follow.","marker":"[69]"},{"why":"Provides the multi-model Greenland projection ensemble that motivates the need for formal uncertainty quantification.","marker":"[31]"},{"why":"Supplies the climate forcing used for surface mass balance and atmospheric temperature boundary conditions.","marker":"[92]"},{"why":"Provides the observed surface velocities from which the basal friction uncertainty model is built.","marker":"[50]"},{"why":"Supplies one of the geothermal heat flux fields used to construct the uncertainty model for the basal thermal boundary condition.","marker":"[90]"}],"fun_headline_variants":["Multifidelity UQ: Ice-sheet forecasts 91x faster, unbiased","Unbiased multifidelity UQ: Greenland melt in days, not years","Multifidelity estimators cut ice-sheet UQ cost 91x","Two-orders speedup: Ice-sheet UQ with unbiased surrogates","91x cheaper ice-sheet UQ: Same accuracy, less compute"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The cost of the pilot runs used to estimate the table of correlations between model predictions is excluded from every speedup calculation, so the advertised savings assume that table is already available for free.","fun_headline_variants_meta":{"raw":{"variants":["Multifidelity UQ: Ice-sheet forecasts 91x faster, unbiased","Unbiased multifidelity UQ: Greenland melt in days, not years","Multifidelity estimators cut ice-sheet UQ cost 91x","Two-orders speedup: Ice-sheet UQ with unbiased surrogates","91x cheaper ice-sheet UQ: Same accuracy, less compute"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000936,"raw_usage":{"total_tokens":4038,"prompt_tokens":1012,"completion_tokens":3026,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":2926}},"tokens_in":628,"tokens_out":3026,"duration_ms":21706,"temperature":1.0,"reasoning_tokens":2926,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:59:45.913390+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recount all speedups with the 32 high-fidelity and 96 low-fidelity pilot runs included in each estimator's budget, and compare the resulting cost with Monte Carlo at the same 1 mm sea-level target; if the multifidelity estimates no longer cost less, the headline speedup claim fails for that accounting.","supporting_citations":[{"cited_title":"Optimal model management for multifidelity Monte Carlo estimation","cited_arxiv_id":null,"evidence_quote":"Supplies the multifidelity Monte Carlo estimator with optimal weights and sample sizes that Section 3.2 implements."},{"cited_title":"Multilevel Monte Carlo methods","cited_arxiv_id":null,"evidence_quote":"Provides the multilevel Monte Carlo telescoping-sum formulation and sample-size balancing used in Section 3.3."},{"cited_title":"On multilevel best linear unbiased estimators","cited_arxiv_id":null,"evidence_quote":"Introduces the multilevel best linear unbiased estimator whose variance the paper compares and extends."},{"cited_title":"Multi-output multilevel best linear unbiased estimators via semidefinite programming","cited_arxiv_id":null,"evidence_quote":"Supplies the semidefinite-programming sample-size optimization that makes the best linear unbiased estimator practical for the 13-model set."},{"cited_title":"Continental scale, high order, high spatial resolution, ice sheet modeling using the Ice Sheet System Model (ISSM)","cited_arxiv_id":null,"evidence_quote":"Is the community ice-sheet code used to run the high-fidelity and surrogate simulations."},{"cited_title":"Experimental protocol for sealevel projections from ISMIP6 standalone ice sheet models","cited_arxiv_id":null,"evidence_quote":"Defines the projection protocol and 2015-2050 setup that the experiments follow."},{"cited_title":"The future sea-level contribution of the Greenland ice sheet: a multi-model ensemble study of ISMIP6","cited_arxiv_id":null,"evidence_quote":"Provides the multi-model Greenland projection ensemble that motivates the need for formal uncertainty quantification."},{"cited_title":"Cnrm-cerfacs cnrm-cm6-1 model output prepared for cmip6 cmip","cited_arxiv_id":null,"evidence_quote":"Supplies the climate forcing used for surface mass balance and atmospheric temperature boundary conditions."},{"cited_title":"Joughin, B","cited_arxiv_id":null,"evidence_quote":"Provides the observed surface velocities from which the basal friction uncertainty model is built."},{"cited_title":"Inferring surface heat flux distributions guided by a global seismic model: particular application to Antarctica","cited_arxiv_id":null,"evidence_quote":"Supplies one of the geothermal heat flux fields used to construct the uncertainty model for the basal thermal boundary condition."}],"review_version":1}