{"id":"a0e8d4ff-de4d-4c68-8099-ddaa6512cda3","arxiv_id":"2501.14982","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"For net-proton C5/C1 and C6/C2 in 0-40% centrality of Au+Au at 11.5 GeV, the CBWC-II event-weighted per-bin ratio method reduces statistical errors to 60-70% of CBWC-I, while error-weighted CBWC-III makes 0-10% collisions negligible.","lead":"A heavy-ion analysis paper compares three ways to correct for volume fluctuations while measuring fifth and sixth order net-proton cumulant ratios. It recommends computing the ratio in each multiplicity bin before averaging, since this cuts statistical errors by 30-40% and avoids an error-weighting scheme that discards the most central collisions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CBWC-II's smaller error is not evidence of accuracy unless the 0-40% target is defined: with centrality-dependent chi ratios, CBWC-I and CBWC-II converge to different population quantities, and the Skellam tests cannot distinguish them.","rationale":"The reader flagged finite-bin ratio bias as the key assumption. That is a related but secondary issue: bias can be checked numerically and appears small in the Skellam averages. The more load-bearing problem is that the target quantity itself is underdefined. Eq. (8) applies to a single thermal system; 0-40% is an aggregate over many systems with different susceptibilities. CBWC-I and CBWC-II estimate different aggregates, and the paper's 'closer to theoretical expectations' claim relies on treating the event-weighted mean of per-bin ratios as the obvious target. This is a conceptual gap that persists even with arbitrarily large statistics, whereas finite-sample bias disappears in the large-sample limit. The UrQMD comparison cannot resolve the gap because UrQMD provides no independent ground truth for the 0-40% ratio. The Skellam simulations provide useful validation of statistical behavior but, by construction, all bins have the same ratio, so they cannot adjudicate between the two estimators' targets. The paper is still valuable as a systematic comparison of CBWC variants and the warning against error-weighted averaging is well supported. My recommendation remains conditional on the authors specifying the intended 0-40% observable and showing that the two estimators' infinite-statistics limits converge, or that the chosen limit is the physically relevant one.","tokens_in":10360,"tokens_out":12143,"duration_ms":125445,"concrete_test":"Use the full 125M-event UrQMD sample to compute the infinite-statistics limits of both estimators in 0-40%: R_II = sum_r n_r (C6_r/C2_r) / sum_r n_r and R_I = (sum_r n_r C6_r) / (sum_r n_r C2_r), and likewise for C5/C1. If R_I and R_II differ by more than the BES-II-level statistical uncertainty, the two methods are estimating different observables, and the accuracy claim is not established until the paper specifies which limit is the desired 0-40% ratio. If the two limits agree within errors, the concern is resolved and the smaller CBWC-II variance is a genuine improvement.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central recommendation presumes there is a single true C_m/C_n for 0-40% centrality that CBWC-II estimates with smaller error. But Eq. (8) only equates C_m/C_n with chi_m/chi_n for a single thermodynamic system. When results from many Refmult3X bins spanning 0-40% are combined, chi_m/chi_n generally varies across bins; UrQMD Fig. 1 itself shows that net-proton ratios in 0-10% differ markedly from those in 10-40%. At infinite statistics, CBWC-I (Eq. 10) converges to a volume-weighted ratio of extensive cumulants, while CBWC-II (Eq. 11) converges to an event-weighted mean of per-bin susceptibility ratios. These are different functionals of the same underlying centrality-dependent distributions, so they need not coincide. The Skellam simulations cannot reveal this because every Refmult3X bin has the same true ratio (unity), forcing both estimators to the same limit. The reported 30-40% error reduction is therefore not, by itself, evidence that CBWC-II is more accurate; it may simply be a smaller-error estimator of a different quantity. The paper needs to specify the intended 0-40% observable; Eq. (8) does not select the event-weighted mean of per-bin ratios over the ratio of volume-corrected cumulants.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares three centrality bin width correction (CBWC) procedures for the net-proton cumulant ratios C5/C1 and C6/C2 in 0-40% centrality in Au+Au collisions at sqrt(s_NN) = 11.5 GeV. Using 125M UrQMD events and Skellam-based Monte Carlo samples with BES-II-like statistics, the authors find that directly averaging per-bin cumulant ratios (CBWC-II) yields statistical errors roughly 60-70% of those from taking the ratio of event-weighted cumulants (CBWC-I), while an error-weighted combination of per-bin results (CBWC-III) makes the 0-10% contribution negligible. The paper recommends CBWC-II and offers a baseline for BES-II analyses.","tokens_in":10646,"tokens_out":9967,"duration_ms":83485,"significance":"If the recommendation is sound, the paper provides a practical, readily implementable prescription for reducing the statistical uncertainty of hyper-order cumulant ratios in a broad centrality bin, with the concrete quantitative claims that the CBWC-II errors for C6/C2 and C5/C1 are 68.8% and 63.7% of the CBWC-I errors, respectively. The use of realistic Refmult3X-based event classification, bootstrap error estimation, and the cross-check between a transport model (UrQMD) and a statistical baseline (Skellam) are strengths. The central issue is that the paper does not define the population quantity that a 0-40% cumulant ratio is intended to estimate, and the simulation evidence does not distinguish between two different functionals when per-bin susceptibility ratios vary.","major_comments":[{"comment":"The paper's central claim that CBWC-II 'aligns more closely' with the theoretical expectation in Eq. (8) is not well-defined, because Eq. (8) applies to a single thermodynamic system. When the true susceptibility ratio chi_m/chi_n varies among Refmult3X bins, as the UrQMD results in Fig. 1 show for 0-10% versus 10-40%, CBWC-I in Eq. (10) and CBWC-II in Eq. (11) converge to different population functionals: a volume-weighted ratio of extensive cumulants versus an event-weighted mean of per-bin susceptibility ratios. The Skellam simulations in Fig. 2 cannot distinguish these functionals because every bin has the same true ratio (unity). The paper should specify the intended 0-40% observable and justify that the event-weighted mean of per-bin ratios is the quantity of theoretical interest; otherwise the smaller statistical error of CBWC-II does not by itself establish that it is a more accurate estimator.","section":"Section 2, Eqs. (10)-(11)"},{"comment":"The accuracy claim for CBWC-II rests on the Skellam simulation, but that simulation only tests the null case of constant per-bin ratios. The paper does not quantify the finite-sample bias of the per-bin ratio estimators (C_m/C_n)_r, which is particularly relevant for C6/C2 in Refmult3X bins with fewer than 0.1M events, as stated in Section 3.2. Although the 30-sample averages are reported to be 'near unity,' the paper does not give the residual bias or a statistical test of bias. Without an explicit bias assessment or a simulation with centrality-dependent true ratios, the observed 60-70% error reduction cannot be interpreted as a reduction in mean squared error.","section":"Section 3.1 and Fig. 2"},{"comment":"The paper states that with widening centrality bin width, variations in the measured values of C6/C2 and C5/C1 are observed between CBWC-I and CBWC-II, but it does not report the magnitude or statistical significance of these differences for the 0-40% bin in the UrQMD model. Without this information, the reader cannot assess whether the 60-70% error reduction is accompanied by a shift in the central value, which is essential for deciding which method is more appropriate for a physics analysis.","section":"Section 3.1 and Fig. 1"}],"minor_comments":[{"comment":"The word 'UrQDM' in the Section 3 heading should be 'UrQMD'.","section":"Section 3 heading"},{"comment":"There are typographical errors in the text: 'C WBC-II' should be 'CBWC-II', and 'the the initial volume fluctuations' should read 'the initial volume fluctuations'.","section":"Section 2"},{"comment":"The sentence 'The formula of CBWC-II method is more equivalent to the theoretical formula as shown in Eq. (11)' is circular, since Eq. (11) is the definition of CBWC-II; the intended comparison is presumably with Eq. (8).","section":"Section 3.1"},{"comment":"The notation 'Refmult3x' and 'Refmult3X' are used inconsistently; please standardize.","section":"Throughout"},{"comment":"Reference [38] appears to duplicate reference [34] with incorrect details; please verify the citation.","section":"References"},{"comment":"The caption says 'the respective error ratios calculated using CBWC-II and CBWC' but should say 'CBWC-II and CBWC-I'.","section":"Fig. 2 caption"},{"comment":"The statement 'the error-weighted average can not be used to calculate these two cumulant ratios in any scenario' is too broad; the demonstration covers a specific Skellam parameter set and one centrality bin. Consider softening to 'in the scenarios considered here'.","section":"Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and addresses a practical question for BES-II analyses. The main revision should focus on the definition of the target 0-40% observable and on validating the accuracy claim for CBWC-II with simulations that allow per-bin susceptibility ratios to vary. These points are fixable with additional analysis, and I do not see evidence of a fundamental error in the simulation setup itself."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a practical analysis-methods paper, not new physics. It does something useful: with BES-II-like statistics, it compares three CBWC variants for net-proton C5/C1 and C6/C2 in 0-40% centrality, shows CBWC-II reduces statistical errors to roughly 60-70% of CBWC-I, and demonstrates that error-weighted CBWC-III makes the 0-10% contribution vanish. Those are real, quantified observations, and the warning against error-weighted averaging is backed by a concrete simulation. The paper deserves referee time.\n\nWhat is genuinely new: the specific comparison of CBWC-I versus CBWC-II for hyperorder cumulant ratios over a wide 0-40% bin with 125M UrQMD events and Skellam ensembles matching BES-II multiplicity and mean multiplicities. The earlier literature discussed CBWC variants, but not this quantified comparison for C5/C1 and C6/C2. No code or data are shipped, but the simulations are simple enough to reproduce and the numbers are consistent with the formulas.\n\nThe soft spots are in proportion. The strongest claim — that CBWC-II \"more precisely measures\" — overreaches because the 0-40% observable is not uniquely defined. Equation (8) applies to a single thermodynamic system. Once you merge Refmult3X bins across 0-10% through 30-40%, CBWC-II converges to an event-weighted mean of per-bin susceptibility ratios, while CBWC-I converges to the ratio of volume-weighted cumulants. If per-bin ratios are centrality-dependent, as UrQMD Fig. 1 itself shows, these are different functionals, and the Skellam tests cannot distinguish them because every bin has true ratio unity, forcing both estimators to the same limit. The paper should state what the intended 0-40% quantity is; otherwise \"smaller error\" is not by itself evidence of accuracy. A test where the true ratio varies across bins, or a bias-quantified comparison of both estimators against a known target, would settle this.\n\nTwo smaller points: per-bin ratio estimators are nonlinear and can be biased at finite occupancy; the paper does not quantify this bias for C5/C1 and C6/C2. And the demonstration that error-weighted averages underestimate, using one Skellam parameter set, is illustrative rather than systematic. These do not sink the main practical message but should be tightened.\n\nBottom line: for heavy-ion fluctuation analysts preparing BES-II analyses, this is a useful, careful reference. It should go to peer review; the referees should push on the target-observable definition and finite-statistics bias before the CBWC-II prescription is adopted as standard.","headline":"A useful, statistically grounded methods paper on CBWC variants for hyperorder cumulants; the recommendation is plausible but the accuracy claim needs a clearer target definition.","tokens_in":11180,"tokens_out":2137,"would_cite":true,"duration_ms":20139,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that the correct CBWC implementation for net-proton cumulant ratios is to compute each ratio inside its multiplicity bin first, and that this cuts statistical errors by roughly a third at the statistics of the next…","keywords":["high-order cumulants","net-proton fluctuations","centrality bin width correction","heavy-ion collisions","QCD phase transition","cumulant ratios","beam energy scan","statistical uncertainty"],"falsifier":"Generate many Monte Carlo samples with the same per-bin multiplicities but with known susceptibility ratios that are not unity, then check whether the event-weighted average of the bin-by-bin ratios reproduces the known ratio on average; a systematic deviation growing as the number of events per bin decreases would falsify the claim that CBWC-II is unbiased. A direct experimental check is to compare CBWC-I and CBWC-II results on real data with the full statistics of the second beam-energy-scan phase: if the CBWC-II values move away from the baseline while CBWC-I values do not, the reduction in error would not correspond to a more accurate measurement.","tokens_in":10182,"feed_emoji":"⚛️","tokens_out":10147,"duration_ms":80407,"temperature":0.7,"pith_summary":"The paper addresses a practical question in heavy-ion data analysis: how the Centrality Bin Width Correction (CBWC) should be implemented when measuring the fifth- and sixth-order cumulant ratios $\\mathrm{C}_5/\\mathrm{C}_1$ and $\\mathrm{C}_6/\\mathrm{C}_2$ of net-proton distributions, observables proposed as probes of the QCD phase transition. It compares three CBWC schemes. The central claim is that computing each cumulant ratio inside its multiplicity bin and then averaging over bins with event-number weights (\"CBWC-II\") removes volume fluctuations in the way the theory intends, and that at the statistics of the second beam-energy-scan phase this reduces statistical errors to roughly 60–70% of those from the ratio-of-weighted-cumulants scheme (\"CBWC-I\"), an effective doubling of the data sample. The paper further finds that an error-weighted combination of 10%-wide centrality bins (\"CBWC-III\") yields smaller errors but discards the 0–10% centrality information, making the 0–40% and 10–40% results identical. Based on a relativistic transport model and difference-of-two-Poisson Monte Carlo samples at 11.5 GeV in Au+Au, it recommends CBWC-II for 0–40% centrality analyses and supplies a baseline for experimental hyperorder-cumulant measurements.","feed_headline":"Ratio-first CBWC cuts net-proton cumulant errors by a third","feed_subtitle":"Computing the ratio bin-by-bin matches the theory of volume cancellation and equals doubling the data sample.","key_machinery":"The central object is the choice among three CBWC averaging formulas. CBWC-I computes $\\mathrm{C}_m$ and $\\mathrm{C}_n$ separately inside each multiplicity bin, weights each cumulant by its event count, and only then divides. CBWC-II computes the ratio $(\\mathrm{C}_m/\\mathrm{C}_n)_r$ in each bin and then event-averages, $\\sum_r \\omega_r (\\mathrm{C}_m/\\mathrm{C}_n)_r$. CBWC-III computes $\\mathrm{C}_m/\\mathrm{C}_n$ in the four 10%-wide centrality bins with event weights and then combines them with inverse-variance weights. The argument is carried by matching the bin-by-bin ratio formula to the susceptibility ratio $\\chi_m/\\chi_n$ after each bin's volume cancels, and by comparing the three estimators in large Monte Carlo samples with a known baseline value of unity.","core_discovery":"On the paper's own terms, the discovery is that the estimator that computes $\\mathrm{C}_m/\\mathrm{C}_n$ bin-by-bin and then event-averages, $\\sum_r \\omega_r (\\mathrm{C}_m/\\mathrm{C}_n)_r$, is the correct CBWC implementation for hyperorder cumulant ratios, while the common estimator that forms $\\sum_r n_r \\mathrm{C}_m^r / \\sum_r n_r \\mathrm{C}_n^r$ retains a residual volume effect and larger statistical error. The bin-by-bin form matches the thermodynamic identity $\\mathrm{C}_m/\\mathrm{C}_n = \\chi_m/\\chi_n$, because each bin's volume cancels before the average is taken. In the model calculations, the 0–40% centrality errors of $\\mathrm{C}_6/\\mathrm{C}_2$ and $\\mathrm{C}_5/\\mathrm{C}_1$ under CBWC-II are 68.8% and 63.7% of the CBWC-I errors, respectively, and thirty independent Monte Carlo samples confirm this 60–70% pattern while keeping the averages consistent with the unit baseline. The paper also establishes that error-weighting the four narrow centrality bins into 0–40% (CBWC-III) makes the 0–10% bin negligible, which is undesirable because the most central collisions carry the strongest phase-transition sensitivity.","pith_inferences":["The variance reduction is a general property of ratio estimators: averaging ratios before dividing suppresses bin-to-bin volume fluctuations, so the same ranking of CBWC schemes should hold for other cumulant ratios such as $\\mathrm{C}_4/\\mathrm{C}_2$ and $\\mathrm{C}_3/\\mathrm{C}_1$, though the magnitude of the gain will depend on the bin occupancy.","If real data reproduce the simulated pattern, adopting CBWC-II could sharpen the energy dependence of $\\mathrm{C}_6/\\mathrm{C}_2$ across the beam-energy scan without any new data taking; the practical cost is only a change in the analysis recipe.","A natural testable extension is to report both CBWC-I and CBWC-II results in the published experimental analysis, so that the bias–variance tradeoff is transparent and theory comparisons are not tied to one estimator choice.","The CBWC-III pathology is a caution for any aggregation scheme that weights by inverse variance: when errors scale with the measured value, inverse-variance weighting can silently drop the physics-rich bin, so centrality windows should be checked for whether each contributing bin actually enters the final number."],"forward_implications":["For 0–40% centrality analyses of $\\mathrm{C}_6/\\mathrm{C}_2$ and $\\mathrm{C}_5/\\mathrm{C}_1$, the paper's recommendation is to use CBWC-II, since it cuts statistical errors to about 60–70% of the CBWC-I values, equivalent to doubling the sample size.","Using error-weighted averages inside 10%-wide bins in central collisions underestimates $\\mathrm{C}_6/\\mathrm{C}_2$, so event-weighting should be used at that stage.","CBWC-III makes the 0–40% and 10–40% results indistinguishable, so any analysis using it forfeits sensitivity to the 0–10% bin where phase-transition signals are most likely.","The model-driven baseline near the difference-of-two-Poisson expectation gives experimental analyses a reference for separating phase-transition effects from the statistical behavior of the CBWC procedure.","Because the 0–10% bin shows significantly different values from the other bins in the transport model, a dedicated high-statistics 0–10% measurement remains necessary despite the appeal of wide-bin error reduction."],"supporting_citations":[{"why":"Defines the centrality bin width correction procedure that the three variants in this paper are built from.","marker":"[33, 34, 35]"},{"why":"Supplies the relativistic transport event generator used for the model comparisons.","marker":"[43, 44]"},{"why":"Establishes the difference-of-two-Poisson distribution as the unit baseline for these cumulant ratios.","marker":"[40, 41, 42]"},{"why":"Sets the centrality-determination multiplicity definition and kinematic cuts that match the second phase of the beam energy scan.","marker":"[45]"},{"why":"Provides the experimental measurements of these hyperorder cumulants that motivate the method comparison.","marker":"[16, 17]"},{"why":"Shows that event-weighted averaging is preferred over error-weighted averaging, a key premise for rejecting the underestimated scheme.","marker":"[37, 38]"},{"why":"Explains why the statistical error of the sixth-order cumulant ratio depends on the measured value, supporting the argument against error-weighted averages.","marker":"[39]"},{"why":"Supplies the bootstrap procedure used to estimate the statistical errors reported in the model comparisons.","marker":"[46, 47]"}],"fun_headline_variants":["Ratio-first CBWC cuts net-proton cumulant errors by a third","Bin-by-bin ratio CBWC matches theory, slashes errors 30%","New CBWC order cuts C6/C2 errors to 69% in heavy-ion data","Hyperorder cumulant ratios: correct CBWC estimator found"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The recommendation assumes that the bin-by-bin ratio estimator is an approximately unbiased estimate of the true susceptibility ratio at the occupancy levels of the second beam-energy-scan phase; ratio estimators are nonlinear and can acquire bias in bins with few events, and the paper does not quantify that bias for $\\mathrm{C}_5/\\mathrm{C}_1$ and $\\mathrm{C}_6/\\mathrm{C}_2$, while also presuming that the transport model and the difference-of-two-Poisson simulation reproduce the statistical behavior of real Au+Au events.","fun_headline_variants_meta":{"raw":{"variants":["Ratio-first CBWC cuts net-proton cumulant errors by a third","Bin-by-bin ratio CBWC matches theory, slashes errors 30%","New CBWC order cuts C6/C2 errors to 69% in heavy-ion data","Hyperorder cumulant ratios: correct CBWC estimator found"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00068,"raw_usage":{"total_tokens":3135,"prompt_tokens":1035,"completion_tokens":2100,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":651,"completion_tokens_details":{"reasoning_tokens":2028}},"tokens_in":651,"tokens_out":2100,"duration_ms":14142,"temperature":1.0,"reasoning_tokens":2028,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:43:29.032125+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate many Monte Carlo samples with the same per-bin multiplicities but with known susceptibility ratios that are not unity, then check whether the event-weighted average of the bin-by-bin ratios reproduces the known ratio on average; a systematic deviation growing as the number of events per bin decreases would falsify the claim that CBWC-II is unbiased. A direct experimental check is to compare CBWC-I and CBWC-II results on real data with the full statistics of the second beam-energy-scan phase: if the CBWC-II values move away from the baseline while CBWC-I values do not, the reduction in error would not correspond to a more accurate measurement.","supporting_citations":[{"cited_title":"Talked at CPOD2024, May 20-14, 2024","cited_arxiv_id":null,"evidence_quote":"Sets the centrality-determination multiplicity definition and kinematic cuts that match the second phase of the beam energy scan."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Explains why the statistical error of the sixth-order cumulant ratio depends on the measured value, supporting the argument against error-weighted averages."}],"review_version":1}