{"id":"f472244f-5311-4263-9cbc-87f9af780c5d","arxiv_id":"2501.08988","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A two-stage simulation-based inference method, using dropout neural networks and normalizing flows, produces fast approximate credibility regions for 3+1 sterile neutrino global fits.","lead":"A new machine learning pipeline draws approximate parameter-fit regions for sterile neutrino experiments, claiming to run about 200 times faster than the standard Feldman-Cousins method. The authors demonstrate it on reactor and source data and suggest the same approach could speed up global fits across particle physics.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The validity of the claimed 200x faster global fit rests on the unvalidated sufficiency assumption in Eq. 2 (p(y|yhat,x)≈p(y|yhat)); the coverage checks in Sec. III C do not validate it and actually show miscalibration, so the speedup has not yet been tied to trustworthy credibility regions.","rationale":"I read the paper in good faith as a proof-of-concept. The strongest claim combines a large runtime speedup with approximate Bayesian inference in a global sterile-neutrino fit. The runtime comparison is transparent and plausible. The load-bearing scientific premise is the conditional-independence/sufficiency assumption in Eq. 2: that the stochastic network output yhat captures all information in x relevant to y. This is not established by the coverage checks, which are necessary but not sufficient; averaged coverage can hide local bias, and the observed deviations in Fig. 4 are evidence of miscalibration. The paper itself flags the arbitrary prior and the lack of a well-defined physical posterior in Sec. IV B, further undercutting the abstract's 'full-fledged Bayesian analysis' wording. None of this makes the approach worthless; it supports a conditional verdict requiring validation of the posterior approximation on held-out points against a high-fidelity reference. The reader's CONDITIONAL verdict is appropriate, and my proposed test would settle whether the concern actually lands, so the verdict should remain UNCHANGED.","tokens_in":14044,"tokens_out":5752,"duration_ms":67052,"concrete_test":"Take 100 parameter points y sampled from the training prior; for each, simulate x with sblmc and compute both the FCMLC posterior via Eq. 2 and a reference posterior using the exact sblmc likelihood with the same prior via MCMC. Compare the 90% highest-posterior-density regions using a 2D Wasserstein distance or a pre-registered containment/coverage metric. If the median distance exceeds a pre-specified tolerance (e.g., 10% of the prior range) or empirical coverage differs by more than 5 percentage points at 90% CL, the sufficiency/prior assumption behind Eq. 2 is falsified and the claimed approximation is not reliable for exclusion curves.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Eq. 2 is formally correct, but the FCMLC posterior is obtained by replacing p(y|yhat,x) with p(y|yhat), estimated from training pairs (y,yhat). This is equivalent to asserting that the dropout-network output yhat is a sufficient statistic for x with respect to y. No direct check of sufficiency is provided. The coverage study in Sec. III C samples y from the training prior and averages over parameter space; good marginal coverage can coexist with locally biased posteriors through overcoverage, and Fig. 4 itself shows deviation from the 45-degree line. Sec. IV B further concedes that the output is not a posterior under a physically motivated prior and that exclusion curves depend strongly on the arbitrary log-space training grid. Thus the abstract's characterization of a \"full-fledged Bayesian analysis\" is not supported for the contours that carry the physics conclusions. The runtime advantage in Table II is plausible, but its scientific value depends on the untested equivalence between FCMLC regions and correct inference regions. The authors are transparent about the proof-of-concept nature of the work, and the qualitative agreement with STEREO, BEST, and the Wilks-based global fit is encouraging; the concern is specifically that the central claim overstates what has been demonstrated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces FCMLC, a simulation-based inference procedure intended to approximate Bayesian posterior distributions for global fits in particle physics. A dropout neural network maps each experimental realization x to a compact prediction y_hat, a KDE converts repeated network calls into p(y_hat|x), a normalizing flow learns p(y|y_hat), and Eq. (2) integrates over y_hat to estimate p(y|x). The method is demonstrated on electron neutrino/antineutrino disappearance data from STEREO, BEST, and a combined global fit, and the authors claim a roughly 200-fold runtime reduction relative to a Feldman-Cousins analysis. The paper is explicitly framed as a proof-of-concept, and it honestly lists several limitations, including imperfect coverage and the training grid acting as an implicit prior.","tokens_in":14400,"tokens_out":3428,"duration_ms":35345,"significance":"If the method's validity were established, it would offer a practical way to perform approximate global fits in regimes where Feldman-Cousins is computationally prohibitive. The qualitative agreement with STEREO's exclusion curve, BEST's allowed region, and the Wilks-based global fit is encouraging and suggests that the architecture can capture relevant structure. The paper is also transparent about its limitations, which is a strength. However, the central claim of approximating a 'full-fledged Bayesian analysis' is not yet supported by the presented validation, primarily because the sufficiency assumption behind Eq. (2) is untested and the coverage checks show miscalibration.","major_comments":[{"comment":"The simplification p(y|y_hat,x) ≈ p(y|y_hat) asserts that the neural-network output y_hat is a sufficient statistic of the data x for the parameters y. This is the load-bearing step of the method, yet no diagnostic is provided to test whether the network output captures all relevant information. The coverage checks in Sec. III C and Fig. 4 do not validate this sufficiency: they sample y from the training prior and can show good marginal coverage even when the conditional posterior is biased, because overcoverage in some regions compensates undercoverage in others. The visible deviation from the 45-degree line in Fig. 4 is consistent with this concern. The abstract's characterization of the result as 'a full-fledged Bayesian analysis' is therefore not supported by the evidence presented.","section":"II B, Eq. (2)"},{"comment":"The authors state that exclusion curves drawn from the estimated posterior depend strongly on the choice of training grid, which acts as an implicit prior. This undermines the interpretation of the drawn regions as Bayesian credibility regions under a physically motivated prior. Moreover, the coverage-check procedure in Sec. II C 2 samples parameters from the same log-uniform grid, so it demonstrates only that the method is self-consistent with its training distribution, not that it produces correct posterior inference for arbitrary physical parameters. The paper should either recalibrate the procedure to yield valid frequentist or Bayesian coverage over the parameter space of interest, or rescope the claims to describe the output as an approximate, prior-dependent surrogate rather than a Bayesian posterior.","section":"IV B"},{"comment":"The coverage miscalibration is acknowledged but only conjecturally attributed to instability of low-credibility contours. Because the paper's value proposition is speed combined with trustworthy inference regions, the miscalibration needs to be quantified and, if possible, mitigated—for example by posterior recalibration or by restricting claims to the high-credibility region where the authors argue the agreement is adequate. Without such analysis, the reader cannot assess whether the 200x speedup is achieved at an acceptable cost in statistical validity.","section":"III C, Fig. 4"},{"comment":"The runtime comparison in Table II is plausible, but the headline factor of ~200 compares a frequentist Feldman-Cousins confidence-interval construction with a procedure that produces approximate posterior regions under an implicit training-grid prior. These are not equivalent inference outputs. The paper should state more explicitly what the speedup is buying and how the FCMLC regions should be interpreted relative to frequentist confidence intervals, especially because the method does not provide a direct frequentist coverage guarantee. This is essential for readers who may use the runtime figure to decide whether FCMLC can replace or supplement Feldman-Cousins analyses.","section":"IV A, Table II"}],"minor_comments":[{"comment":"There are typos such as 'frequentest' for 'frequentist' and 'uncetainties' for 'uncertainties'; the paper would benefit from a careful proofread.","section":"I, Introduction"},{"comment":"The sentence 'We alter replaced the conditional KDE with a normalizing flow' contains a grammatical error; it should read 'We therefore replaced the conditional KDE with a normalizing flow'.","section":"II B"},{"comment":"The description of the coverage-check prior states that parameters are sampled 'uniformly in the logarithm of each parameter'; it would be helpful to state explicitly that this matches the training-grid prior used in Sec. II A, since that prior is central to the interpretation of the coverage results.","section":"II C 2"},{"comment":"The sentence 'we MC samples were generated from grid points uniformly separated in log10 Ue4 and log10 ∆m2_41 space' is awkwardly phrased and should be rewritten for clarity.","section":"IV B"},{"comment":"The notation 'R² = 0' in the discussion of Fig. 2 is not defined; a brief explanation of what R² refers to in this context would improve readability.","section":"Fig. 2"}],"recommendation":"major_revision","confidential_remarks":"The paper has a mismatch between the abstract's 'full-fledged Bayesian analysis' claim and the limitations stated in Sec. IV B, where the authors concede that the procedure does not compute a posterior under a well-defined physical prior. This inconsistency should be resolved before publication. Additionally, there is no explicit data or code availability statement, which is worth noting for a paper that presents a computational method; the reproducibility of the 510,000-hour Feldman-Cousins comparison and the FCMLC training details would be aided by making the code and generated datasets available."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a credible, transparent proof-of-concept that a two-stage SBI stack (dropout NN -> KDE -> normalizing flow) can generate approximate posterior-like regions for a multi-experiment sterile neutrino disappearance fit about 200x faster than the reference Feldman-Cousins implementation. The qualitative agreement with published STEREO, BEST, and the Wilks-based global fit is encouraging. But the paper's headline phrase, 'full-fledged Bayesian analysis,' is not supported. The posterior depends on an arbitrary log-space training grid, and the key conditional-independence assumption in Eq. 2 is asserted, not validated.\n\nWhat's actually new: the components are established (Gal & Ghahramani dropout, normalizing flows, law-of-total-probability two-stage estimation), but the integrated application to a multi-experiment global fit in neutrino physics is novel. The authors also deserve credit for candor. In Sec. IV B they say outright that FCMLC does not compute a posterior under a well-defined prior, that exclusion curves 'depend strongly' on the training grid, and that coverage is imperfect. That honesty makes the paper more useful than a slicker one with the same gaps.\n\nThe soft spots are real and tied to the central claim. Eq. 2 replaces p(y|yhat,x) with p(y|yhat), which is effectively an assertion that the network output is a sufficient statistic for the data with respect to the parameters. No direct check of that sufficiency is provided. The coverage checks in Fig. 4 sample from the training prior and still show visible miscalibration, especially at lower confidence levels — exactly the regime where exclusion curves are drawn. The authors wave this off by noting that neutrino physics cares about >90% CL, but a statement like that doesn't make the regions trustworthy. The runtime comparison is also somewhat soft: the 510k-hour figure comes from an unoptimized MCMC-based FC implementation, and the 2.5k-hour total includes the shared MC generation, so '200x' is not apples-to-apples. The qualitative speedup is still real, but the factor is roughly stated. Finally, no code or data are released, which makes independent checks harder than they should be.\n\nWho this is for: neutrino physicists who want a fast way to scan parameter space before committing to a full Feldman-Cousins fit, and SBI practitioners interested in multi-experiment applications. The paper deserves serious peer review: it is a legitimate proof-of-concept with a real runtime gain and an honest limitations section. A referee should push for a validation of the sufficiency assumption (e.g., a likelihood-based diagnostic or a held-out simulation check) and for softening the language from 'full-fledged Bayesian analysis' to what the method actually delivers: a fast, approximate, empirically-calibrated surrogate. I would not cite the current version in my own work in the next twelve months, but I would read a revised version carefully.","headline":"Honest and useful SBI proof-of-concept for sterile neutrino global fits, but the 'full-fledged Bayesian analysis' claim outruns what is validated; the speedup is real, the credibility regions are not yet.","tokens_in":14891,"tokens_out":3310,"would_cite":false,"duration_ms":30878,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A machine-learning pipeline can run sterile neutrino global fits about 200 times faster, the authors report.","keywords":["simulation-based inference","sterile neutrinos","global fits","Feldman-Cousins method","normalizing flows","Bayesian neural networks","credibility regions","neutrino oscillation"],"falsifier":"Compare two mock data sets that the trained network maps to the same predicted parameters but that were generated from different true model points; if the FCMLC posteriors for the two data sets differ materially, the conditional-independence approximation $p(y|\\hat{y},x) \\approx p(y|\\hat{y})$ fails. A more direct version would bin test samples by the predicted parameters and check whether the distribution of true parameters is independent of the original data within each bin; a systematic dependence would falsify the method's core approximation.","tokens_in":13841,"feed_emoji":"⚛️","tokens_out":11598,"duration_ms":103566,"temperature":0.7,"pith_summary":"For small-signal particle physics analyses, Wilks' theorem can fail, and the standard fix, the Feldman-Cousins procedure, becomes computationally intractable when many experiments are combined. This paper claims that a simulation-based inference pipeline, dubbed FCMLC, can approximate a full Bayesian analysis about 200 times faster, reducing a sterile neutrino global fit from roughly 510,000 to 2,500 computing hours. The pipeline replaces the high-dimensional data-to-parameter problem with a two-step density estimate: a stochastic neural network predicts the oscillation parameters, and a normalizing flow learns the conditional distribution of true parameters given those predictions. The authors demonstrate the method on a joint analysis of electron neutrino and antineutrino disappearance experiments under a 3+1 sterile neutrino model, reproducing the qualitative features of published single-experiment limits and of a Wilks-based global fit. If the claim holds, it makes full global fits feasible for small-signal searches where Feldman-Cousins is too expensive.","feed_headline":"Machine-learning pipeline cuts sterile neutrino fit time 200-fold","feed_subtitle":"Simulation-based inference approximates global-fit results while cutting runtime from 510,000 to 2,500 hours.","key_machinery":"The load-bearing object is the posterior identity $p(y|x) \\approx \\int p(y|\\hat{y})\\, p(\\hat{y}|x)\\, d\\hat{y}$, which follows from the law of total probability together with the approximation that the neural network's predicted parameters $\\hat{y}$ screen off the data $x$ from the true parameters $y$. The machinery has three parts: a dropout-at-inference neural network that turns one experimental data vector into many randomized parameter predictions, a kernel density estimate that converts those predictions into $p(\\hat{y}|x)$, and a normalizing flow, an invertible neural transformation of a base distribution into a flexible conditional density, that supplies $p(y|\\hat{y})$. Numerical integration over a 50 by 50 grid in $\\hat{y}$ and thresholding of the resulting $p(y|x)$ builds the credibility regions used for fits and exclusions.","core_discovery":"The paper's central claim is that the posterior density over the true oscillation parameters $y$ given experimental data $x$ can be written, under the conditional-independence approximation $p(y|\\hat{y},x) \\approx p(y|\\hat{y})$, as $p(y|x) \\approx \\int p(y|\\hat{y})\\, p(\\hat{y}|x)\\, d\\hat{y}$. Here $\\hat{y}$ are the parameters predicted by a dropout-at-inference neural network, $p(\\hat{y}|x)$ is estimated by a kernel density estimate over repeated network calls, and $p(y|\\hat{y})$ is learned by a normalizing flow. Credibility regions are obtained by numerically solving for the level set of this approximate posterior that contains the desired probability mass. The authors show that for STEREO-only, BEST-only, and combined electron-flavor disappearance fits, FCMLC produces contours that qualitatively match the published STEREO exclusion, the BEST allowed region, and a Wilks-based global fit, including the gallium-anomaly region, while reducing the total runtime by a factor of about 200 relative to a Feldman-Cousins treatment.","pith_inferences":["The conditional-independence assumption is effectively a sufficiency claim about the learned network output. A direct test would be to compare FCMLC posteriors for two data vectors that map to the same predicted parameters; any substantial difference would indicate that $x$ carries information beyond $\\hat{y}$, and the coverage mismatch seen in the paper's Fig. 4 may already be a symptom of this.","Because the training grid is uniform in log-space and the authors note that exclusion curves depend strongly on that choice, the credibility regions should be read as prior-dependent summaries rather than frequentist confidence intervals until a physically motivated prior is introduced.","The current demonstration is restricted to two oscillation parameters; extending to all 3+1 channels would test whether the normalizing flow's density estimate and the conditional-independence approximation survive in higher dimension, and could be done with modest changes to the pipeline.","One could calibrate FCMLC coverages empirically, adjusting credibility thresholds until observed coverage matches nominal coverage, to obtain frequentist-style limits at a fraction of the cost of a full Feldman-Cousins calculation."],"forward_implications":["Sterile neutrino global fits on electron neutrino and antineutrino disappearance data become practical, taking about 2,500 hours instead of about 510,000 hours for the same realization generation.","Single-experiment fits, and the extension from one experiment to a global combination, add little runtime once the neural network and normalizing flow are trained.","FCMLC reproduces the qualitative shapes of the published STEREO and BEST results and of the Wilks-based global fit, including the parameter region associated with the gallium anomaly.","Because every component is modular, the same recipe can be applied to other small-signal global fits where Wilks' theorem fails and Feldman-Cousins is too costly.","The fast posterior can serve as a first pass to localize parameter space before running a higher-fidelity Feldman-Cousins analysis."],"supporting_citations":[{"why":"Defines the Feldman-Cousins confidence-belts procedure whose computational cost is the baseline FCMLC aims to beat.","marker":"[14]"},{"why":"Documents failures of Wilks' theorem in small-signal mock trials, establishing why the expensive alternative is needed.","marker":"[13]"},{"why":"Supplies the dropout-as-Bayesian-approximation technique used to make the first-stage neural network stochastic.","marker":"[26]"},{"why":"Introduces normalizing flows for density estimation, the basis of the conditional density module.","marker":"[28]"},{"why":"Provides the neural spline flow architecture used to model $p(y|\\hat{y})$.","marker":"[29]"},{"why":"Provides the frequentist fitting framework used to generate realizations and the Wilks-based global fit for comparison.","marker":"[30]"},{"why":"Provides the previous 3+1 global fit and the gallium-anomaly best-fit point that the FCMLC credibility regions are compared against.","marker":"[20]"},{"why":"Supplies STEREO data and the published exclusion curve used as a single-experiment reference.","marker":"[1]"},{"why":"Supplies BEST data and the published allowed region used as a single-experiment reference.","marker":"[8]"}],"fun_headline_variants":["ML inference makes sterile neutrino fits 200x faster","Neural network speeds sterile neutrino global fits 200-fold","Simulation-based inference cuts neutrino fit time 200x","200x faster sterile neutrino global fits via ML","Drop-in ML accelerates neutrino global fits by 200x"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes the network's predicted parameters capture all the information in the data that matters for the true parameters, so that once the prediction is known, the original data tells you nothing more about the physics.","fun_headline_variants_meta":{"raw":{"variants":["ML inference makes sterile neutrino fits 200x faster","Neural network speeds sterile neutrino global fits 200-fold","Simulation-based inference cuts neutrino fit time 200x","200x faster sterile neutrino global fits via ML","Drop-in ML accelerates neutrino global fits by 200x"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001102,"raw_usage":{"total_tokens":4600,"prompt_tokens":950,"completion_tokens":3650,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":566,"completion_tokens_details":{"reasoning_tokens":3572}},"tokens_in":566,"tokens_out":3650,"duration_ms":25861,"temperature":1.0,"reasoning_tokens":3572,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:12:19.323195+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare two mock data sets that the trained network maps to the same predicted parameters but that were generated from different true model points; if the FCMLC posteriors for the two data sets differ materially, the conditional-independence approximation $p(y|\\hat{y},x) \\approx p(y|\\hat{y})$ fails. A more direct version would bin test samples by the predicted parameters and check whether the distribution of true parameters is independent of the original data within each bin; a systematic dependence would falsify the method's core approximation.","supporting_citations":[{"cited_title":"org/10.3103/S1068335623120151","cited_arxiv_id":null,"evidence_quote":"Defines the Feldman-Cousins confidence-belts procedure whose computational cost is the baseline FCMLC aims to beat."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents failures of Wilks' theorem in small-signal mock trials, establishing why the expensive alternative is needed."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the dropout-as-Bayesian-approximation technique used to make the first-stage neural network stochastic."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the neural spline flow architecture used to model $p(y|\\hat{y})$."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the frequentist fitting framework used to generate realizations and the Wilks-based global fit for comparison."},{"cited_title":"However, for global fits, it’s crucial that our networks can quantify predictive uncer- tainties meaningfully","cited_arxiv_id":null,"evidence_quote":"Supplies STEREO data and the published exclusion curve used as a single-experiment reference."},{"cited_title":"II A), use sblmc to simulate experimental data from said model parameters, and then use fcmlc to build an estimate of the posterior distribution from which CRs can be drawn","cited_arxiv_id":null,"evidence_quote":"Supplies BEST data and the published allowed region used as a single-experiment reference."}],"review_version":1}