{"id":"cf9eca9b-7ac2-4b3a-a911-f55f94a2270a","arxiv_id":"2507.18824","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Simulation-based inference trained on synthetic scattering data yields rho(770) pole estimates closer to reference values than chi-squared minimization in the tested misspecification cases.","lead":"Using a neural-network simulation approach, the authors estimate the rho(770) resonance pole from pion-pion scattering data. When the fitted model does not perfectly describe the data, their method gives pole positions closer to reference values than standard chi-squared fitting in the tested cases.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The chi2-anchored training-data filter in Sec. III C induces a prior that already favors the SBI result in the n=2 cases, so the real-data advantage may be an artifact of the filter rather than of SBI.","rationale":"The reader's conditional verdict already rests on this assumption, and my analysis agrees: the training filter is the weakest load-bearing point. I do not think the concern alone warrants rejection, because the toy examples in Sec. II B are not subject to the filter and provide independent evidence that SBI can outperform chi2 under the two constructed misspecifications; the n=3 rows of Table III also suggest the filter is not always dominant. However, the n=2 rows show the filter-induced prior mean is already closer to SBI than to chi2 by 4-5 MeV, while the claimed SBI advantage in those cases is 9-18 MeV; the paper's one-sentence rebuttal does not separate this contribution. A controlled sensitivity run would settle it. The paper is transparent about the limitation, which makes the issue a fixable methodological gap rather than a sign of misreporting. Hence I keep the reader's CONDITIONAL verdict.","tokens_in":19918,"tokens_out":5358,"duration_ms":58266,"concrete_test":"Re-run the n=2 Estabrooks and Protopopescu SBI analyses with the Sec. III C training filter centered on shifted chi2 poles, e.g., shift Re E* by -20, -10, +10, +20 MeV and Im E* by -10, -5, +5, +10 MeV in a crossed design, keeping all other settings fixed. Compute the resulting SBI pole positions; if they move by more than about 5 MeV (the typical SBI-vs-chi2 separation in Table II/III), the filter is a dominant bias and the real-data comparison is not a clean test. If the output is stable under such shifts, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing weakness is the data-dependent training filter in Sec. III C. Pseudodata are retained only if their chi2-fitted pole lies in the square centered on the chi2 pole of the actual data, so the effective SBI prior is itself a function of the chi2 result the method is compared against. The paper explicitly acknowledges that 'anchoring the training data to the chi2 can effect the final predictions,' and its own Table III shows this is not negligible: for both n=2 cases the average training pole is closer to the SBI pole than to the chi2 pole (9.1 vs 13.2 MeV and 8.3 vs 10.6 MeV), a 4-5 MeV offset in the same direction as the SBI-vs-chi2 differences. The n=3 cases are less affected (average training pole far from both), which is why the toy examples, which have no such filter, still support the general method. But the four-case real-data claim is at least half-supported by n=2, where the apparent SBI advantage may be a property of the filter-induced prior rather than of SBI's handling of misspecification. The paper's rebuttal that the average training pole is 'roughly equidistant' from the two methods does not quantify the size of the induced bias relative to the effect being claimed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes the use of neural-network-based Simulation Based Inference (SBI) as an alternative to chi-squared minimization for extracting resonance pole positions when the fitted model is misspecified. The method is first demonstrated on two toy models with known ground truth, where SBI is claimed to recover the generating parameters substantially better than chi2 minimization. The method is then applied to pi-pi scattering phase-shift data from two experiments (Protopopescu and Estabrooks) using two K-matrix models with two and three subtractions, giving four real-data cases. In all four cases the SBI pole is reported to be closer to the PDG reference values than the chi2 pole. A classifier network is also used to select the number of subtractions, and a combined fit is performed. The paper includes appendices documenting tests of the neural network architecture and of the convergence of the reported uncertainties.","tokens_in":20161,"tokens_out":7284,"duration_ms":71549,"significance":"If the central claim is correct, the paper provides a useful proof of concept that SBI can be more robust than chi2 minimization under model misspecification, with direct relevance to hadron spectroscopy where data ambiguities are common. The paper is clearly written and gives a reproducible description of the pipeline, including the pseudodata generation, the neural-network architecture, and empirical checks in Appendices A and B. The toy examples are valuable because the ground truth is known and the observed differences between SBI and chi2 are large. However, the real-data claim is currently confounded by a data-dependent training filter whose influence is only partially quantified, and the comparison to PDG values is not accompanied by a statistical test of significance. These issues affect the load-bearing conclusion of the paper and need to be addressed before the claim can be accepted.","major_comments":[{"comment":"The chi2-anchored retention window is a data-dependent prior that can bias the SBI prediction toward the training distribution. The paper's own Table III shows that for the two n=2 cases, the average training pole is closer to the SBI pole than to the chi2 pole: 9.1 vs 13.2 MeV for Estabrooks and 8.3 vs 10.6 MeV for Protopopescu, whereas the SBI-to-chi2 distances are 18.5 and 9.5 MeV, respectively. The text in Sec. III C states that the average training pole is \"roughly equidistant\" from the two methods, but the differences are 4.1 and 2.3 MeV, which are 22% and 24% of the corresponding SBI-chi2 separations. This is not negligible, and it is exactly in the direction of the claimed SBI advantage. The statement in Sec. III C that the retention step \"ensures that, if the SBI method makes a more accurate prediction than the chi2 method, it is not because it is trained on data that is closer to known position\" is therefore misleading: the training distribution is closer to the SBI prediction than to the chi2 prediction. A sensitivity analysis with retention windows centered on a range of positions spanning the chi2 and SBI poles, or an otherwise reweighted prior, is needed to establish that the n=2 real-data results are not an artifact of the filter.","section":"Sec. III C, Table III"},{"comment":"The claim that SBI is \"more accurate\" or \"as close or closer\" is based on point estimates without a statistical measure of the difference. The toy examples in Sec. II B are single realizations; no repeated-experiment or coverage analysis is reported that would show how often SBI beats chi2 across random draws of the misspecified data. For the real-data cases, the differences between SBI and chi2 are sometimes comparable to the reported one-sigma uncertainties. For instance, in the Estabrooks n=3 case of Table II the real parts differ by 769.6 - 766.8 = 2.8 MeV, while the SBI uncertainty on that quantity is 2.9 MeV. A bootstrap of the difference, or a paired test over the N=100 network realizations, should be reported before the conclusion of a systematically more accurate method is drawn.","section":"Sec. II B and Sec. III A"}],"minor_comments":[{"comment":"The sentence \"SBI is shown to make a more robust predictions\" is ungrammatical; it should read \"SBI is shown to make more robust predictions\" or \"a more robust prediction\".","section":"Abstract"},{"comment":"The text states that the SBI width is closer to the true value \"by more than an order of magnitude\" than the chi2 width. From Table I, the distance to the true value is 174.5 MeV for SBI (293.5 vs 119.0) and 363.6 MeV for chi2 (482.6 vs 119.0), a factor of about 2.1. This statement should be corrected or substantiated.","section":"Sec. II B, text after Table I"},{"comment":"\"The training data is given to a neural network\" should be \"The training data are given\", since \"data\" is plural.","section":"Sec. III C"},{"comment":"The caption refers to \"boxes\" but the figure is a flow diagram; the boxes in the middle and bottom rows are not clearly labeled in the text. Please add labels or refer to the components explicitly in the caption.","section":"Fig. 1"},{"comment":"The uncertainty notation \"0.06482(00095)\" is nonstandard and confusing; use \"0.06482(95)\" or \"0.06482 \\pm 0.00095\".","section":"Eq. (4)"},{"comment":"The SBI columns list the fitted K-matrix parameters a_i without uncertainties, while the pole-position uncertainties are derived from the spread of the N network outputs. Reporting the parameter uncertainties would aid reproducibility and comparison with the chi2 results.","section":"Table II"}],"recommendation":"major_revision","confidential_remarks":"The paper's own Table III directly undercuts the \"roughly equidistant\" claim in Sec. III C: for both n=2 cases the average training pole is several MeV closer to the SBI pole than to the chi2 pole, and this shift is a non-negligible fraction of the claimed SBI-vs-chi2 separation. The authors explicitly acknowledge that anchoring can affect predictions, but they do not quantify the induced bias relative to the effect size. I recommend requiring a sensitivity analysis of the retention window before considering the paper for publication. Also, the \"order of magnitude\" statement for the toy-model width should be corrected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is a genuine proof-of-concept: neural-network SBI can give more accurate rho(770) pole positions than chi2 minimization when the model is misspecified. The toy examples with known ground truth are convincing, and the application to two data sets and two K-matrix models is a legitimate, well-scoped case study. The new element is the explicit misspecification focus—earlier ML pole-extraction papers didn't frame it that way, and the robust-SBI literature hadn't done a hadronic application.\n\nWhat the paper does well: the method is described clearly, the pseudodata generation is transparent, and the authors repeatedly flag their own limitations. They say the confidence regions don't account for bias, and they note that the four cases are not independent. The classifier network for choosing the number of subtractions is a nice addition.\n\nThe soft spot is the training-data filter in Sec III C. Pseudodata are kept only if their chi2-fitted pole falls in a square around the chi2 pole of the actual data, so the effective prior is a function of the chi2 result the method is compared against. The stress-test note is on point: Table III shows the n=2 training poles are 2-4 MeV closer to the SBI result than to the chi2 result, in the same direction as the claimed gain. The paper calls this 'roughly equidistant,' but the numbers don't fully support that. This doesn't sink the paper—the n=3 cases are much less affected and the toys have no such filter—but it means the 'all four cases' claim is half-supported by a possible confound. I'd want either a direct test of the filter's influence or a retention rule not anchored to the chi2 pole.\n\nMinor issues: the real-data 'truth' is the PDG reference, not a known pole; each toy is a single realization; and there's no code or data artifact. None of these are deal-breakers, but they cap the certainty.\n\nThe paper deserves a serious referee. The target audience is hadronic spectroscopists and SBI practitioners; both would get something out of it. Recommendation: send it out, and in revision require a quantitative test of the filter bias and a corrected interpretation of Table III.","headline":"A credible SBI-vs-chi2 proof-of-concept whose real-data claim is weakened by a chi2-anchored training filter that the authors underplay.","tokens_in":20775,"tokens_out":3569,"would_cite":true,"duration_ms":34180,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Simulation-based inference yields more accurate rho(770) pole positions than chi-squared minimization under model misspecification.","keywords":["simulation based inference","model misspecification","resonance pole position","rho(770)","pi-pi scattering","neural network","K-matrix","chi-squared minimization"],"falsifier":"Retrain the SBI network with the retention window centered on a different point—for example, on the reference pole instead of the chi-squared pole—and compare the predicted pole position. If the prediction follows the center, the reported advantage is an anchoring artifact; if it does not move, the advantage is robust. The same comparison across several independent data sets would quantify how general the claim is.","tokens_in":19664,"feed_emoji":"⚛️","tokens_out":8368,"duration_ms":80963,"temperature":0.7,"pith_summary":"This paper claims that when the model being fitted does not actually describe the data, a neural-network-based Simulation Based Inference (SBI) pipeline can estimate the rho(770) resonance pole more accurately than traditional chi-squared minimization. The demonstration starts with two toy models—one with intentionally corrupted points and one with non-Gaussian noise—where the true parameters are known, and then moves to actual pi-pi scattering phase-shift data fitted with a K-matrix amplitude using two and three subtractions. In all four model-and-data combinations, the SBI pole position is as close as or closer to reference values than the pole from chi-squared minimization, even though the chi-squared curve sits closer to the data points themselves. The broader reason to care is that many hadron spectroscopy analyses rely on data with documented ambiguities, so a method that resists misspecification could change how resonance parameters are extracted. The paper explicitly limits its claim: in near-ideal settings chi-squared minimization remains the better approach.","feed_headline":"Neural nets beat chi-squared for rho(770) poles","feed_subtitle":"When pi-pi scattering models misspecify the data, simulation-based inference returns resonance parameters closer to reference values.","key_machinery":"The central mechanism is the SBI training pipeline: sample $10^6$ random K-matrix parameters from broad priors, generate Gaussian pseudodata at the experimental energy points with the same uncertainties as the real data, keep only pseudodata whose fitted pole falls inside a square centered on the chi-squared pole from the actual data, and train a four-hidden-layer neural network to predict the generating parameters from the pseudodata curves. The training is repeated $N=100$ times and the average prediction is quoted with uncertainty $\\sigma/\\sqrt{N}$. The physical object being estimated is the resonance pole $E^* = M - i\\Gamma/2$ obtained by analytic continuation of the K-matrix amplitude with $n$ subtractions. The comparison baseline is standard chi-squared minimization with bootstrap uncertainties on the same models and data.","core_discovery":"On the paper's own terms, the discovery is a proof of concept: under model misspecification, SBI delivers resonance parameters closer to the true or reference values than chi-squared minimization does. In the toy examples, the SBI mass and width are far closer to the generating values, especially the width in the outlier example. In the real-data application, four combinations of two data sets and two subtraction numbers are considered, and in every case the SBI pole $E^* = M - i\\Gamma/2$ lies closer to the reference values than the chi-squared pole; for the three-subtraction combined fit the data show clear misspecification ($\\chi^2/\\mathrm{dof}=3.1$, $p=0.17\\times10^{-3}$), yet SBI still lands closer to the reference values. The paper interprets SBI as an approximate Bayesian inference in which the random parameter draw serves as a prior and the pseudodata encode the likelihood, so the network approximates $P(\\vec a \\mid \\vec y)$ rather than the single best-fitting curve.","pith_inferences":["A decisive test beyond the paper would be to rerun the training with the pole-retention window centered on the reference pole instead of the chi-squared pole; if the SBI answer moves with the window, part of the reported advantage is an artifact.","The same procedure could be tried on resonances with notoriously inconsistent data, such as the $\\Lambda(1405)$ two-pole structure, where the underlying model uncertainty is larger than here.","If the advantage survives de-anchoring, one practical consequence is that legacy resonance parameters extracted by chi-squared fits from old data sets may need rechecking with misspecification-robust methods.","The dependence of SBI on the chosen parameter prior ranges can be measured directly by widening or shifting the uniform priors and observing how the final pole prediction changes."],"forward_implications":["Because SBI does not need to identify which data points are flawed, it can be used when the source of misspecification is unknown, as long as pseudodata can be generated from a candidate model.","The classifier part of the pipeline can pick the most probable model (here, three subtractions) from pseudodata before the final parameter extraction, so model choice and parameter estimation can be joined in one workflow.","The same pipeline transfers to other resonances whose data are ambiguous, such as the $a_1(1260)$ and $\\omega(782)$ cases mentioned in the paper, without changing the core procedure.","In near-ideal fits the chi-squared method is still expected to be better, so SBI is a complementary tool for misspecified data rather than a replacement for standard practice.","The combined three-subtraction SBI fit provides a candidate set of $\\rho(770)$ pole parameters closer to reference values than the corresponding chi-squared fit, which matters for downstream analyses that use $\\pi\\pi$ scattering input."],"supporting_citations":[{"why":"Introduces the simulation-based inference concept that the paper adapts.","marker":"[1]"},{"why":"Establishes approximate Bayesian computation, the statistical family the pipeline instantiates.","marker":"[4]"},{"why":"Documents difficulties of SBI under misspecification, motivating the robust version used here.","marker":"[12]"},{"why":"Provides a misspecification-robust SBI variant cited as context for the chosen implementation.","marker":"[14]"},{"why":"Supplies one of the two pi-pi phase-shift data sets used in the application.","marker":"[17]"},{"why":"Supplies the other pi-pi phase-shift data set used in the application.","marker":"[18]"},{"why":"Provides the K-matrix-like scattering amplitude with a variable number of subtractions used by both fitting methods.","marker":"[44]"},{"why":"Supplies the reference pole values that define which method is more accurate.","marker":"[98]"}],"fun_headline_variants":["SBI improves rho(770) poles under model misspecification","Neural simulation inference bests chi-squared for rho poles","Model misspecification? SBI still finds the rho pole","Chi-squared misleads; SBI finds rho pole accurately"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that restricting the SBI training set to pseudodata whose fitted poles lie in a square centered on the chi-squared pole does not itself move the network's prediction toward the SBI answer; the paper's own check shows the average training pole is closer to the SBI result than to the chi-squared result in the two-subtraction cases, so the premise is only partially supported.","fun_headline_variants_meta":{"raw":{"variants":["SBI improves rho(770) poles under model misspecification","Neural simulation inference bests chi-squared for rho poles","Model misspecification? SBI still finds the rho pole","Chi-squared misleads; SBI finds rho pole accurately"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00025,"raw_usage":{"total_tokens":1595,"prompt_tokens":1028,"completion_tokens":567,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":644,"completion_tokens_details":{"reasoning_tokens":492}},"tokens_in":644,"tokens_out":567,"duration_ms":5615,"temperature":1.0,"reasoning_tokens":492,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:07:55.204653+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the SBI network with the retention window centered on a different point—for example, on the reference pole instead of the chi-squared pole—and compare the predicted pole position. If the prediction follows the center, the reported advantage is an anchoring artifact; if it does not move, the advantage is robust. The same comparison across several independent data sets would quantify how general the claim is.","supporting_citations":[{"cited_title":"Predicting the mpemba effect using machine learning,","cited_arxiv_id":null,"evidence_quote":"Supplies the other pi-pi phase-shift data set used in the application."}],"review_version":1}