{"id":"4774ac32-2bec-40f6-9b26-fdb60a546e0a","arxiv_id":"2411.10223","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Using higher moments of air-shower depth distributions, the authors project that the full Pierre Auger dataset could reject the EPOS and Sibyll hadronic models at high confidence.","lead":"This paper checks whether the hadronic models used to simulate cosmic-ray air showers can explain the measured shower-depth distributions from the Pierre Auger Observatory. It finds that current public data do not rule them out, but that including higher statistical moments would let the full Auger dataset clearly distinguish between the models.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Rejection CLs from Eq. (4) rely on an unjustified chi-square_n assumption after fitting w* to the same data; calibration is needed before the f≥5 claims stand.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing issue: Eq. (4) is treated as chi-square_n despite w* being fitted to the same moments. This is the single most important problem because the paper's headline quantitative results—CL ≥ 95% at f = 5 for EPOS and Sibyll with n = 5—are derived directly from that distributional claim. If the calibration is wrong, the specific CL values are not meaningful, and the assertion that the existing full Auger dataset would suffice to discriminate hadronic models is unsupported. The concern is concrete and testable with standard pseudo-experiment calibration. I considered other potential issues, such as whether the finite 2000-shower simulation sample places a floor on the projected uncertainty at large f; while this is also a real limitation, it affects only the f-scaling of the simulated covariance and is secondary to the fundamental miscalibration of the test statistic. The paper itself contains honest caveats about the projection (fixed central values, conservative treatment of detector resolution), and the qualitative tension in higher moments is visible in the pulls of Fig. 5. These are genuine merits. Nevertheless, the quantitative CLs cannot be accepted as calibrated without either accounting for the fitted parameters in the degrees of freedom or performing the calibration check. Since the reader's verdict of CONDITIONAL already captures this requirement (the claim should be accepted only after calibration), my stress-test does not move the verdict; it confirms that the condition is necessary. The empirical calibration described above would settle the matter: if the distribution of Eq. (4) matches chi-square_n, the concern is resolved and the reported CLs stand; if not, the CLs must be revised or the claims softened. I therefore maintain the reader's CONDITIONAL verdict as UNCHANGED.","tokens_in":10022,"tokens_out":7844,"duration_ms":80352,"concrete_test":"Run a Monte Carlo calibration under the null hypothesis. For a chosen model (e.g., EPOS) and a chosen n (e.g., n = 5), take the best-fit composition w* reported at f = 1 as the true composition. Generate a large number (≥10^4) of pseudo-datasets, each with N = f·N_PAOD events sampled from the simulated Xmax distributions for that composition, and apply the exact same pipeline: bootstrap moments, fit w* via nested-sampling likelihood maximization, and compute the statistic in Eq. (4). Compare the empirical distribution of this statistic to the claimed chi-square_n distribution, for f = 1, 5, 10, 20. If the empirical 90th and 95th percentiles lie significantly below the chi-square_n thresholds, the reported rejection CLs are overestimated and must be recalibrated. Report the empirical mean and variance of the statistic to quantify the effective number of degrees of freedom.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—that EPOS and Sibyll reach CL ≥ 95% at f = 5 when n = 5—rests on the assertion in Sec. II.C that the statistic in Eq. (4) follows a chi-square distribution with n degrees of freedom. This assertion is not justified because w* (the 26-component composition, with 25 independent fractions after the sum constraint) is fitted to the very same measured moments that enter Eq. (4). Consequently, the simulated mean µ_s(w*) is correlated with the measured mean µ_m, and the covariance of their difference is not Σ_s + Σ_m as stated. Fitting parameters to the data typically reduces the effective number of degrees of freedom of a quadratic goodness-of-fit statistic; with more fitted fractions than moments, the post-fit discrepancy is dominated by what the model cannot absorb, not by a chi-square_n fluctuation. The paper neither subtracts the number of fitted parameters nor calibrates the distribution via pseudo-experiments. The reported rejection CLs are therefore uncalibrated: the right-tail probabilities derived from chi-square_n do not correspond to the true probability of observing such a discrepancy under the null hypothesis. Since the headline result (CL ≥ 95% already at f = 5 for both models) is expressed in those CLs, the claim is quantitatively unsupported until the calibration is established. The qualitative tension visible in the pulls of Fig. 5 may survive recalibration, but the specific CL values and the statement that the existing full Auger dataset would suffice to discriminate models do not follow from the analysis as presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper tests the compatibility of four hadronic interaction models (EPOS, Sibyll, QGSJetII-04, QGSJet01) with Pierre Auger Open Data by comparing the first n central moments of the Xmax distribution. For each model, a 26-component mass composition w* is fitted via a bootstrap-based likelihood, and a rejection confidence level is computed from the post-fit difference between simulated and measured moments using Eq. (4), which is asserted to follow a chi-square distribution with n degrees of freedom. The main quantitative result is that for EPOS and Sibyll, adding the fourth and fifth moments (n = 4, 5) makes the rejection CL grow with the statistical multiplier f, reaching CL >= 95% already at f = 5. The paper also studies energy-spectrum reweighting, concluding the effect is negligible in a narrow energy bin, and reports rejection results for the QGSJet models in an appendix.","tokens_in":10369,"tokens_out":4121,"duration_ms":44727,"significance":"If the statistical calibration were sound, this would be a valuable use of public Auger data to discriminate among current hadronic interaction models with a modest amount of new methodology. The paper has clear strengths: it uses public data and CORSIKA simulations, includes all 26 primaries up to iron, propagates systematic and statistical uncertainties through a bootstrap procedure, and makes an explicit projection to larger statistics. The energy-reweighting study in Sect. III is a useful systematic check. However, the central rejection CLs rest on an uncalibrated chi-square assumption, so the quantitative claims—in particular 'CL >= 95% already at f = 5'—are not yet supported. The qualitative tension visible in the pulls may survive recalibration, but the specific confidence levels and the projection to the full Auger dataset require a corrected or empirically calibrated test statistic.","major_comments":[{"comment":"The statement that z_s - z_m follows N(mu_s - mu_m, Sigma_s + Sigma_m) and that Eq. (4) follows a chi-square distribution with n degrees of freedom is not justified, because mu_s = mu_s(w*) and Sigma_s = Sigma_s(w*) are evaluated at the composition w* obtained by maximizing the same log-likelihood against the measured moments z_m. The composition vector has 26 components (25 independent fractions), so for n = 3, 4, 5 the number of fitted composition parameters exceeds the number of moments. Fitting parameters to the same data generally reduces the effective number of degrees of freedom of a quadratic goodness-of-fit statistic and introduces correlation between mu_s and z_m, so the covariance of their difference is not simply Sigma_s + Sigma_m. The paper neither subtracts the number of fitted parameters nor calibrates the distribution of Eq. (4) with pseudo-experiments. Consequently, the right-tail probabilities reported in Figs. 3 and 4—including the headline 'CL >= 95% already at f = 5 for both models'—are uncalibrated and do not currently have the stated frequentist meaning. A calibration procedure that repeats the full composition fit on Monte Carlo pseudo-data is needed before these quantitative claims can be accepted.","section":"II.C, Eq. (4)"},{"comment":"The projection to larger statistics by the factor f assumes that the central values of the measured moments remain fixed while only their covariance shrinks. Under the null hypothesis that a hadronic model is correct, a new realization of a larger dataset would have different central values, and the composition fit would absorb part of those fluctuations. The current procedure therefore computes an in-sample goodness-of-fit rather than an out-of-sample prediction. The paper's own caveat in Sec. IV ('this does not mean that these models will in fact be incompatible once more data is analyzed') is appropriate, but it is in tension with the abstract and Sect. III, where the f = 5 result is presented as evidence that the existing full PA dataset would suffice to reject these models. A proper treatment would either calibrate the post-fit statistic under the null or present the f-dependence only as a sensitivity projection, with the CL interpretation deferred to a calibrated analysis.","section":"III, Figs. 3 and 4"},{"comment":"The appendix's conclusions that QGSJetII-04 and QGSJet01 are already in strong tension with data for n = 1 or n = 2 moments rely on the same uncalibrated chi-square assumption as Eq. (4). In particular, the high CL values close to unity in Fig. 8 are computed without accounting for the fact that the composition w* is fitted to the same moments that enter the statistic. The df problem is even more pronounced for n = 1, where a single moment is compared after fitting 25 independent composition fractions. These appendix claims therefore inherit the calibration issue and should be either re-derived with a calibrated statistic or explicitly labeled as provisional.","section":"Appendix A, Figs. 6 and 8"}],"minor_comments":[{"comment":"The text uses 'P A Open Data' in several places; this should be 'PA Open Data' for consistency.","section":"II.A"},{"comment":"In Eq. (2), the denominator appears as a sum over u_i with a superscript 'l' in the typeset version; this should be simply the sum of u_i, matching the definition in the text.","section":"II.B, Eq. (2)"},{"comment":"In the sentence 'At f = 1, that is the actual PAOD, we have CL ~ 20%', the phrase 'the actual PAOD' is imprecise; use 'the current PAOD subsample' to avoid implying the full dataset.","section":"III"},{"comment":"The statement that new model implementations (EPOS4, Sibyll*, QGSJetIII) 'are expected to be consistent' with the versions used here is an unsupported assertion; it would be helpful to cite a quantitative comparison or to soften the wording.","section":"Footnote 2"},{"comment":"The legend in Fig. 5 is hard to parse because the labels 'f = 1' and 'f = 20' are placed on a single line without separating the line styles for each panel; a per-panel legend would improve readability.","section":"Fig. 5"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe useful part here is the projection: adding the fourth and fifth Xmax moments to the fit makes EPOS and Sibyll look increasingly bad against the Auger fluorescence data once you scale the statistics by f around 5-10. If that survives recalibration, it is a real step toward the muon puzzle. The paper uses the public Auger open data, simulates 26 primaries with CORSIKA, includes energy reweighting and bootstrap errors, and the qualitative pull plots in Fig. 5 show a genuine tension in the higher moments. The authors are also honest in the conclusions that they are projecting sensitivity, not claiming the models are already ruled out.\n\nThe soft spot is Eq. (4). They fit the 26-component composition w* to the same measured moments that then go into the chi-square statistic, and they treat the difference between the simulated and measured moment vectors as a multivariate normal with covariance Sigma_s + Sigma_m and n degrees of freedom. That is not right. Because w* is a function of the data, the post-fit residuals are correlated with the measured moments and biased toward zero; the effective degrees of freedom are far below n, and with 25 independent fractions fitting 3-5 moments, the n=3 consistency is nearly a tautology. The CL values in Figs. 3 and 4 are therefore not calibrated. This is not a footnote issue; it is the load-bearing claim of the paper.\n\nThe fix is straightforward in principle: run pseudo-experiments under the null to calibrate the statistic, or at least subtract an effective number of fitted parameters. The authors do neither. The qualitative tension visible in the pulls may survive any reasonable calibration, but the specific statement that CL >= 95% at f = 5 for EPOS and Sibyll does not follow from the analysis as presented.\n\nWho this is for: people working on UHECR mass composition and hadronic interaction models. The method extension is a legitimate follow-up to the authors' prior work, and the honest caveats in the text help. It deserves a serious referee, not a desk rejection, because the qualitative claim is interesting and the flaw is correctable. My own verdict is conditional: I would not cite the quantitative rejection levels until the statistic is calibrated, but I would read a revised version carefully.","headline":"A promising higher-moment sensitivity study for hadronic models, but the quoted rejection CLs rely on an uncalibrated chi-square that ignores the fitted composition.","tokens_in":10866,"tokens_out":2949,"would_cite":false,"duration_ms":30044,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper asks whether including the fourth and fifth central moments of the shower-depth-maximum (Xmax) distribution can discriminate among leading hadronic interaction models.","keywords":["ultra-high energy cosmic rays","hadronic interaction models","Xmax distribution","central moments","cosmic ray composition","Pierre Auger Observatory","bootstrap","model discrimination"],"falsifier":"A Monte Carlo calibration: simulate many pseudo-datasets from the best-fit composition, refit the composition and evaluate the same test statistic each time; if its 95th percentile lies below the $\\chi^2_n$ 95th percentile, the reported rejection confidence levels for EPOS and Sibyll are overstated.","tokens_in":9846,"feed_emoji":"🔭","tokens_out":5718,"duration_ms":46303,"temperature":0.7,"pith_summary":"This paper asks whether the leading models of high-energy hadronic interactions can survive a stricter test against cosmic-ray data. The authors compare the first few central moments of the depth of shower maximum (Xmax) distribution measured by the Pierre Auger Observatory with predictions from the EPOS, Sibyll, QGSJetII-04, and QGSJet01 models, after fitting the unknown mix of primary nuclei up to iron. They find that with only the first three moments all models pass, but once the fourth and fifth moments are included, EPOS and Sibyll are rejected at 90–95% confidence once the dataset is scaled up by a factor of five to ten. The result matters because it identifies higher Xmax moments as a sharp discriminator between hadronic models, and suggests the existing full Auger dataset has enough statistical power to distinguish them.","feed_headline":"Higher shower-depth moments could rule out two hadronic models","feed_subtitle":"At five times the current statistics, EPOS and Sibyll hit 95% rejection confidence against Auger data.","key_machinery":"The machinery is the comparison of the first $n$ central moments of the Xmax distribution, $z_1, \\dots, z_n$ (mean, variance, skewness, kurtosis, and beyond), between simulated predictions and measured data. The mass composition is a 26-component vector $w$ of primary fractions (proton through iron), fitted by maximizing a log-likelihood built from bootstrap-estimated multivariate normal distributions of the moments; the rejection confidence level is then computed from a $\\chi^2$ statistic comparing the best-fit moment vector to the measured one.","core_discovery":"The paper claims that the first three central moments of the Xmax distribution are not enough to challenge the leading hadronic models, but the fourth and fifth moments make the predictions of EPOS and Sibyll statistically incompatible with Pierre Auger data once the statistics are scaled upward. Concretely, including five moments yields a rejection confidence level above 95% at a statistics multiplier of f = 5 for both models, and above 90% at f ≈ 10 with four moments. The paper stresses that these are projections based on fixed central values of the measured moments, so they demonstrate sensitivity rather than a definitive incompatibility, but they indicate that the full Auger dataset would have sufficient discriminating power.","pith_inferences":["A proper calibration of the test statistic, e.g. by parametric bootstrap, could change the claimed confidence levels; this is the most direct way to test whether the f = 5 rejection survives.","The same moment-based comparison could be turned into a tuning target for next-generation hadronic models: matching the fourth and fifth moments of Xmax may be a stronger constraint than matching only the mean and variance.","If higher moments of Xmax prove to be this sensitive, other shower observables with similarly non-Gaussian tails might offer additional discriminators for hadronic models.","The scaling procedure assumes the central values of the measured moments stay fixed as f grows; any energy-dependent systematic that shifts those central values with statistics would change the projection."],"forward_implications":["With the first three Xmax moments, EPOS and Sibyll fit the Pierre Auger open data well even when the sample is extrapolated to twenty times its size.","Adding the fourth moment makes the rejection confidence grow with statistics, passing roughly 90% near f = 10.","Adding the fifth moment pushes the rejection confidence above 95% already at f = 5 for both EPOS and Sibyll.","QGSJetII-04 and QGSJet01 are already disfavored using only the first one or two moments.","If the measured moment values hold as statistics grow, the full Auger dataset would be enough to discriminate among the four models."],"supporting_citations":[{"why":"Establishes the moment-based composition inference method that this paper extends to higher moments.","marker":"[12]"},{"why":"Introduces the statistics-multiplier f projection used to extrapolate to larger datasets.","marker":"[13]"},{"why":"The Pierre Auger 2021 Open Data release, the measured dataset analyzed here.","marker":"[15]"},{"why":"Provides the systematic uncertainty treatment and detector-smearing procedure applied to the simulated Xmax distributions.","marker":"[16]"},{"why":"CORSIKA, the shower simulation package whose hadronic models are being tested.","marker":"[17]"},{"why":"The EPOS hadronic interaction model, one of the two main case studies.","marker":"[18]"},{"why":"The Sibyll hadronic interaction model, the other main case study.","marker":"[19]"},{"why":"The QGSJetII-04 model, rejected with the first moment in the appendix.","marker":"[20]"},{"why":"The QGSJet01 model, rejected with the second moment in the appendix.","marker":"[21]"},{"why":"The bootstrap method used to turn finite event samples into multivariate normal moment distributions.","marker":"[23]"}],"fun_headline_variants":["Higher moments of shower depth may rule out two hadronic models","Fifth moment of shower depth could rule out EPOS and Sibyll","Higher Xmax moments could challenge hadronic models with more data","With 5x data, higher shower moments reject EPOS and Sibyll","Fourth and fifth shower moments could veto two hadronic models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that its test statistic follows a $\\chi^2$ distribution with $n$ degrees of freedom, even though the 26 composition fractions are fitted to the same measured moments that are then tested; with more fitted parameters than moments, the post-fit discrepancy is likely smaller than the assumed distribution, so the reported confidence levels are not calibrated.","fun_headline_variants_meta":{"raw":{"variants":["Higher moments of shower depth may rule out two hadronic models","Fifth moment of shower depth could rule out EPOS and Sibyll","Higher Xmax moments could challenge hadronic models with more data","With 5x data, higher shower moments reject EPOS and Sibyll","Fourth and fifth shower moments could veto two hadronic models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000737,"raw_usage":{"total_tokens":3215,"prompt_tokens":789,"completion_tokens":2426,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":405,"completion_tokens_details":{"reasoning_tokens":2337}},"tokens_in":405,"tokens_out":2426,"duration_ms":14351,"temperature":1.0,"reasoning_tokens":2337,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:50:00.325391+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A Monte Carlo calibration: simulate many pseudo-datasets from the best-fit composition, refit the composition and evaluate the same test statistic each time; if its 95th percentile lies below the $\\chi^2_n$ 95th percentile, the reported rejection confidence levels for EPOS and Sibyll are overstated.","supporting_citations":[{"cited_title":"On the model uncertainties for the predicted maximum depth of extensive air showers","cited_arxiv_id":"2409.05501","evidence_quote":"The Pierre Auger 2021 Open Data release, the measured dataset analyzed here."},{"cited_title":"Learning the Composition of Ultra High Energy Cosmic Rays","cited_arxiv_id":"2212.04760","evidence_quote":"CORSIKA, the shower simulation package whose hadronic models are being tested."},{"cited_title":"Improving the Composition of Ultra High Energy Cosmic Rays with Ground Detector Data","cited_arxiv_id":"2304.11197","evidence_quote":"The EPOS hadronic interaction model, one of the two main case studies."},{"cited_title":"Resolution of (Heavy) Primaries in Ultra High Energy Cosmic Rays","cited_arxiv_id":"2409.06841","evidence_quote":"The Sibyll hadronic interaction model, the other main case study."},{"cited_title":"Pierre auger observatory 2021 open data,","cited_arxiv_id":null,"evidence_quote":"The QGSJetII-04 model, rejected with the first moment in the appendix."}],"review_version":1}