{"id":"4581d794-3412-4ba3-9dce-f63d4dfe17d1","arxiv_id":"2501.00150","paper_version":3,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of uncertainty quantification for quasi-Monte Carlo that recommends Student's t intervals from at least 10 randomized replicates and identifies near-symmetry of RQMC errors as a promising but unproven basis for confidence intervals.","lead":"This paper surveys how to estimate the error of quasi-Monte Carlo numerical integration, where randomized replicates can give confidence intervals. It highlights a recent empirical surprise: some randomized QMC estimates have nearly symmetric error distributions, so ordinary t-intervals may work without a central limit theorem or a consistent variance estimate.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified.","rationale":"The reader's verdict ACCEPT is sound. The paper's 'strongest claim' is not a new theorem but a synthesis of existing empirical and theoretical results. The weakest point identified by the reader - generalization of the near-symmetry phenomenon - is exactly the caveat the author flags in Sections 5 and 9. Because the survey explicitly attributes the recommendation to [61], describes the limited theoretical explanations in [85], and labels the symmetry mechanism as needing more study, there is no internal inconsistency or overclaim that would justify changing the verdict. A held-out replication of the coverage study is a worthwhile verification of the generalization risk, but its absence is a research direction, not a defect in the survey.","tokens_in":27009,"tokens_out":10909,"duration_ms":108336,"concrete_test":"Re-run the [61] coverage protocol on a held-out grid of integrands and RQMC methods (for example, Owen-scrambled Faure nets, randomly shifted rank-1 lattices, and non-smooth or high-dimensional integrands) with R=10, using 1000 fresh replications per case; if the standard Student-t interval falls below 94% coverage in a nontrivial fraction of the held-out cases, the survey's practical recommendation would need to be narrowed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"No load-bearing objection identified. The paper is a survey, and its central claims are appropriately scoped: the R>=10 Student-t recommendation is presented as an empirical finding from [61], the theoretical support in [85] is explicitly limited to one scrambling method, and Section 9 states that the near-symmetry mechanism 'clearly needs more study.' The main residual risk is that the empirical coverage result may not generalize beyond the five RQMC methods and six integrands tested; however, the paper does not claim otherwise, so this is a stated limitation rather than an internal flaw. Weighing the explicit caveats in Sections 5 and 9 does not change the descriptive accuracy of the survey.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This proceedings survey reviews methods for quantifying the error of quasi-Monte Carlo estimates. It distinguishes certificates, confidence intervals, and asymptotic intervals; covers the Koksma-Hlawka and weighted-space bounds, bracketing and NNLD/NPLD certificates, RQMC and the recent empirical finding that Student's t intervals from R≥10 independent replicates have reliable coverage; and surveys the Warnock-Halton quasi-standard error, GAIL, and new directions (unbiased MCMC, normalizing flows, median-of-means, and R growing with n). The central practical recommendation is the R≥10 rule, with the near-symmetry of RQMC error distributions identified as the key open theoretical mechanism.","tokens_in":27094,"tokens_out":14871,"duration_ms":136053,"significance":"If taken as a survey, the paper is valuable: it organizes a fragmented literature, gives the first accessible account of the RQMC confidence-interval simulation results of [61] and the skewness/kurtosis theory of [84,85], and makes a concrete, falsifiable recommendation. Its strength is careful scoping: Section 5 explicitly states that the empirical findings have not been established for other RQMC methods, and Section 9 states that the symmetry mechanism clearly needs more study. The paper also credits the relevant literature and points out known pitfalls (e.g., the Warnock-Halton QSE failure). No machine-checked proofs are supplied, but the survey does not claim them; its claims are traceable to published sources.","major_comments":[],"minor_comments":[{"comment":"The statement 'Having κ_n→∞ completely rules out a CLT for \\hat μ_n from matrix scrambling with a digital shift' is stronger than what a diverging kurtosis alone implies; a sequence of distributions can have diverging fourth moments and still converge in distribution to a Gaussian. If a stronger obstruction is proved in [84] or [85], it should be cited here; otherwise the wording should be weakened to say that the diverging kurtosis defeats the moment- and variance-estimation-based routes to a CLT confidence interval.","section":"Section 5 (kurtosis of matrix scrambling)"},{"comment":"The phrase 'γ_n=O(n^ε) for any ε>0, so it is almost O(1)' is not strictly correct, since such a sequence may be unbounded (e.g., log n); 'sub-polynomial' would be more accurate.","section":"Section 5 (skewness of random linear scrambling)"},{"comment":"The sentence 'the critical event has probability Ω(d/n^2)' appears inconsistent with the one-dimensional event probability Ω(1/n); for fixed d it would give a fourth-moment contribution Ω(n^{-6}), which does not produce the claimed diverging kurtosis. Please clarify whether the intended rate is Ω(d/n) or explain the extra factor 1/n.","section":"Section 5 (critical event for d>1)"},{"comment":"There are a few typos: 'random random vector' in Section 4, 'erroneosly' in Section 5, and 'ANOV A' in Section 1.","section":"Throughout"}],"recommendation":"accept","confidential_remarks":"The survey draws heavily on the author's own prior work, but the cited papers are published and the survey is transparent about reliance on them. The main residual risk is that the R≥10 recommendation is empirical, but the manuscript explicitly states this limitation. Given the paper's scope as a proceedings survey, I support acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a well-scoped proceedings survey, not a research paper. It has no new theorem of its own—Theorem 1 is explicitly credited to [39], and the main empirical claims come from the simulation study in [61] and the theory in [84,85]. But the paper earns its place by giving a clear map of error estimation for RQMC and stating, with unusual honesty, what is known and what is not.\n\nWhat it does well: the recommendation to use R >= 10 independent replications with a Student's t interval is genuinely actionable, and the paper connects it to Hall's Edgeworth expansions in a way that makes the empirical finding plausible rather than magical. The discussion of why large kurtosis helps the standard t interval and hurts the bootstrap-t is good. The survey also does a service by laying out the gap between certificates, conservative bounds, and asymptotic intervals, and by explicitly flagging the near-symmetry phenomenon as something that clearly needs more study.\n\nSoft spots: the central near-symmetry claim is empirical and narrow. It is observed for five RQMC methods and six integrands, and the only partial theoretical explanation is for random linear scrambling with a digital shift. The paper does not claim more than that, but readers should not walk away thinking this is a theorem. Generalization to other methods is an open question. Also, the paper leans heavily on the author's own prior work; that is self-citation, but the cited results are published and peer-reviewed, so it is not circular. The control-variate suggestion in Section 4 is a suggestion, not a tested method, and the paper says so.\n\nThere is not much to object to. The limitations are in the text, not hidden. I would send this to peer review; a proceedings audience will get real value from it, and a survey of this type does not need a new theorem to deserve a serious referee.","headline":"A careful, well-scoped survey of RQMC error estimation with a practical R>=10 t-interval recommendation; no new theorem, but honest about what is known and a solid contribution to the proceedings.","tokens_in":27588,"tokens_out":2013,"would_cite":true,"duration_ms":20399,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["65D30","65C05","62F25"],"pacs":[],"model":"deepseek-v4-flash","headline":"Randomized quasi-Monte Carlo estimates can get reliable confidence intervals without a central limit theorem, thanks to a near-symmetry that appears as the sample size grows.","keywords":["randomized quasi-Monte Carlo","confidence intervals","Student's t interval","skewness","kurtosis","digital nets","median of means","scrambling"],"falsifier":"Repeat the [61] protocol with R=10 on an RQMC method outside the five tested, such as a random-start Halton sequence or a higher-order scrambled net, using an integrand whose estimator distribution is markedly skewed; if the standard t interval covers μ in fewer than roughly 94% of 1000 repetitions, the recipe fails to generalize. Alternatively, compute γ_n for random linear scrambling on a smooth integrand with a nonzero one-dimensional effect and look for a case where |γ_n| grows faster than O(n^ε), which would refute the theoretical explanation in [85].","tokens_in":26783,"feed_emoji":"🎲","tokens_out":8426,"duration_ms":78067,"temperature":0.7,"pith_summary":"Quasi-Monte Carlo can integrate high-dimensional functions far more accurately than plain Monte Carlo, but its very accuracy makes the error hard to estimate: rigorous bounds are usually uncomputable or far too wide. This survey asks what can still be said about the error of a QMC estimate, and it organizes the available answers into certificates, confidence intervals, and asymptotic intervals. The central finding is an empirical surprise documented in [61] and partially explained in [85]: for several randomized QMC methods the distribution of the estimator turns nearly symmetric as n grows, even when its kurtosis diverges and no central limit theorem holds. Because of that symmetry, the standard Student-t interval based on R >= 10 independent replications gives reliable 95% coverage, without needing a consistent estimate of the estimator variance. A sympathetic reader should care because this yields a simple, practical recipe for trustworthy error bars on high-accuracy QMC computations and identifies exactly where the theory is still open.","feed_headline":"Ten replicates give reliable error bars for quasi-Monte Carlo","feed_subtitle":"A survey finds standard Student-t intervals hit 95% coverage for several randomized QMC methods, even without a CLT.","key_machinery":"The machinery is the distribution of the RQMC estimator μ̂_n, summarized by its skewness γ_n and kurtosis κ_n, together with the Student-t statistic built from R independent replicates. Edgeworth expansions for two-sided coverage error show that the standard interval's error is governed by terms like κ_n/R and $γ_n^{2}$/R, so a method with small skewness but large kurtosis can still have near-nominal two-sided coverage, and heavy-tailed replicate distributions tend to make the t interval slightly conservative. For random linear scrambling of base-2 digital nets, [84] shows there is an event of probability Ω(1/n) with squared error Ω(1/$n^{2}$) that drives κ_n to infinity, while [85] shows γ_n = O(n^ε), which is why the distribution can be non-Gaussian yet symmetric. The rare-outlier structure also explains why the sample variance s̃^2 often underestimates $σ_n^{2}$: with small R the outliers are usually unseen, and symmetry keeps the resulting interval reliable rather than misleading.","core_discovery":"The paper's claim, stated on its own terms, is that uncertainty quantification for QMC is not hopeless: it is a matter of choosing the right tradeoff between accuracy and estimability. With a fixed budget of N=nR function evaluations, larger n gives a better estimate of the integral while larger R gives a better estimate of the error, and RQMC with small R is the regime where both goals can be met. The striking discovery is that some RQMC estimators have distributions that become nearly symmetric as n grows without becoming Gaussian: skewness stays at O(n^ε) or smaller while kurtosis can diverge to infinity. This near-symmetry is enough to make the ordinary Student-t confidence interval from R >= 10 independent replicates attain close to nominal 95% coverage, despite the fact that the sample variance of the replicates can badly underestimate the true variance and no CLT applies. The paper treats this as a surprise that needs more study, and it recommends R >= 10 independent replications with the standard t interval as the current best-supported recipe for RQMC uncertainty quantification.","pith_inferences":["If near-symmetry is the true mechanism, one could deliberately design RQMC randomizations, for example symmetric shifts or paired reflections, to force the estimator distribution symmetric and extend the t-interval guarantee beyond the five methods tested.","The same rare-outlier structure that makes kurtosis diverge is exactly what median-of-means is built to survive; combining a symmetric heavy-tailed estimator with a median or trimmed-mean aggregation could yield both super-polynomial accuracy and a computable confidence interval, an extension not developed in the paper.","The coverage evidence is for two-sided 95% intervals; Edgeworth expansions show one-sided coverage errors do not benefit from the same cancellation, so one-sided or 99% intervals may need larger R, a testable implication the paper does not pursue.","The NNLD/NPLD certificate idea suggests a deterministic analogue: point sets with symmetric discrepancy structure might yield guaranteed two-sided error bounds for completely monotone integrands without any randomization."],"forward_implications":["A practitioner who wants a 95% confidence interval for an RQMC estimate can average R >= 10 independent replications and use the usual Student-t interval, without estimating σ_n^2 or relying on a CLT.","For a fixed budget N=nR, allocating essentially all effort to n leaves too few replicates for UQ; the tradeoff analysis supports keeping R at least 10 even at some cost in raw accuracy.","For matrix-scrambled Sobol' nets with a digital shift, the diverging kurtosis rules out CLT-based confidence intervals, so the near-symmetric, heavy-tailed regime is the correct asymptotic target for that family.","Bootstrap percentile intervals performed badly (1689 failures out of 2400 cases) and bootstrap-t was still worse than the standard t interval, so the standard interval is the strongest current recommendation.","Certificates do exist for special integrand classes, such as bracketing for convex or monotone functions and NNLD/NPLD bounds for completely monotone functions, but they are conservative and suffer a dimension effect; confidence intervals are the practical default."],"supporting_citations":[{"why":"This simulation study of five RQMC methods and six integrands at R in {5,10,20,30} is the empirical basis for the recommendation, since the standard t interval failed only 3 of 2400 cases.","marker":"[61]"},{"why":"It gives the partial theoretical explanation that the skewness of the RQMC estimator under random linear scrambling is O(n^ε), supporting the observed near-symmetry.","marker":"[85]"},{"why":"It proves the kurtosis of the estimator diverges as n grows for analytic integrands under matrix scrambling, establishing that no CLT can hold and motivating the symmetry-based explanation.","marker":"[84]"},{"why":"It provides the CLT for scrambled (t,m,d)-nets with t=0, the benchmark that other RQMC methods fail to reach and the contrast against which symmetry becomes relevant.","marker":"[67]"},{"why":"It reports the small-sample simulation comparison of confidence interval methods whose qualitative pattern, including bootstrap-t failures and the lognormal exception, informs the RQMC findings.","marker":"[75]"},{"why":"It gives the Edgeworth expansions for two-sided coverage error in terms of skewness and kurtosis, the theoretical lens used to explain why the standard t interval survives large kurtosis.","marker":"[45]"},{"why":"It documents the long-known tendency of heavy-tailed distributions to make Student-t intervals conservative, the mechanism invoked to explain the coverage behavior.","marker":"[17]"},{"why":"It supplies the CLT conditions and Lyapunov-type bounds for R growing with n, defining the boundary of the regime where asymptotic intervals are valid.","marker":"[72]"}],"fun_headline_variants":["Ten replicates give reliable error bars for quasi-Monte Carlo","Quasi-Monte Carlo error bars work with just ten replicates","RQMC error estimation: R>=10 yields accurate t-intervals","No CLT needed: ten replicates suffice for QMC error bars"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the near-symmetry seen in the simulation study of five RQMC methods and six integrands, and partially explained for random linear scrambling, holds broadly enough across other RQMC methods and integrands to justify the 'R >= 10 with Student-t' recipe; the paper itself says in Sections 5 and 9 that this has not been established for the other methods and clearly needs more study.","fun_headline_variants_meta":{"raw":{"variants":["Ten replicates give reliable error bars for quasi-Monte Carlo","Quasi-Monte Carlo error bars work with just ten replicates","RQMC error estimation: R>=10 yields accurate t-intervals","No CLT needed: ten replicates suffice for QMC error bars"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000862,"raw_usage":{"total_tokens":3695,"prompt_tokens":856,"completion_tokens":2839,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":472,"completion_tokens_details":{"reasoning_tokens":2767}},"tokens_in":472,"tokens_out":2839,"duration_ms":20144,"temperature":1.0,"reasoning_tokens":2767,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:57:53.660141+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the [61] protocol with R=10 on an RQMC method outside the five tested, such as a random-start Halton sequence or a higher-order scrambled net, using an integrand whose estimator distribution is markedly skewed; if the standard t interval covers μ in fewer than roughly 94% of 1000 repetitions, the recipe fails to generalize. Alternatively, compute γ_n for random linear scrambling on a smooth integrand with a nonzero one-dimensional effect and look for a case where |γ_n| grows faster than O(n^ε), which would refute the theoretical explanation in [85].","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This simulation study of five RQMC methods and six integrands at R in {5,10,20,30} is the empirical basis for the recommendation, since the standard t interval failed only 3 of 2400 cases."},{"cited_title":"Skewness of a randomized quasi-Monte Carlo estimate","cited_arxiv_id":"2405.06136","evidence_quote":"It gives the partial theoretical explanation that the skewness of the RQMC estimator under random linear scrambling is O(n^ε), supporting the observed near-symmetry."},{"cited_title":"Mathematics of Computation 92(340), 805–837 (2023)","cited_arxiv_id":null,"evidence_quote":"It proves the kurtosis of the estimator diverges as n grows for analytic integrands under matrix scrambling, establishing that no CLT can hold and motivating the symmetry-based explanation."},{"cited_title":"Annals of Statistics 31(4), 1282–1324 (2003)","cited_arxiv_id":null,"evidence_quote":"It provides the CLT for scrambled (t,m,d)-nets with t=0, the benchmark that other RQMC methods fail to reach and the contrast against which symmetry becomes relevant."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It reports the small-sample simulation comparison of confidence interval methods whose qualitative pattern, including bootstrap-t failures and the lognormal exception, informs the RQMC findings."},{"cited_title":"The Annals of Statistics 16(3), 927–953 (1988)","cited_arxiv_id":null,"evidence_quote":"It gives the Edgeworth expansions for two-sided coverage error in terms of skewness and kurtosis, the theoretical lens used to explain why the standard t interval survives large kurtosis."},{"cited_title":"ACM Transactions on Modeling and Computer Simulation 34(3), 1–38 (2024)","cited_arxiv_id":null,"evidence_quote":"It supplies the CLT conditions and Lyapunov-type bounds for R growing with n, defining the boundary of the regime where asymptotic intervals are valid."}],"review_version":1}