{"id":"7c4ca9d6-2019-4f21-b9fc-766251f19205","arxiv_id":"2608.02164","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"The ENUBET Demonstrator, a 1.7 m iron-scintillator calorimeter prototype, achieves the electron energy resolution (≈15%/√E stochastic term) and e/π/μ discrimination required for monitored neutrino beams, though SiPM saturation introduces non-linearities.","lead":"A full-size prototype of ENUBET's instrumented decay-tunnel calorimeter — iron absorbers with scintillator tiles read out by optical fibers and silicon photomultipliers — was built and tested at CERN in three beam campaigns. The prototype meets the electron energy resolution and particle-identification requirements needed to monitor neutrino production at the source, with photosensor saturation flagged as a remaining limitation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simulation 'validation' depends on per-particle energy-scale and contamination parameters fitted to the same data; agreement is a closure test, not an independent check of the ENUBET full simulation.","rationale":"The reader's weakest_assumption identifies the same core issue: the data-MC agreement used to claim validation of the full simulation is achieved only after multiple parameters are adjusted to the data. My analysis of the text confirms this concern is load-bearing. The per-particle energy-scale corrections, especially the 18% reduction for pions, are not derived from an independent physics model; they are fitted to the total-energy data. The contamination fractions are also fitted. The saturation model is admittedly imperfect and overcorrects the highest-deposit channel, yet the energy-resolution comparison averages MC with and without saturation to bracket the data. Under these conditions, the agreement in the shower profiles and total-energy distributions does not independently validate the full simulation—it demonstrates that a flexible simulation can be tuned to match the test-beam data. This does not undermine the direct performance claims (energy resolution, light yield, cross-talk), which are measured quantities. But the abstract and conclusions extend the claim to the full ENUBET simulation and its associated 1% flux-systematics projection, and that extension is not supported by this analysis. A conditional verdict—keeping the measured performance claims but requiring the simulation-validation wording to be aligned with the tunable comparison—is appropriate. I agree with the reader's assessment and do not see grounds to move to REJECT, since the measured central quantities are plausible and the limitations are explicitly disclosed. The proposed concrete test would settle whether the pion-scale correction is benign or concealing a hadronic-modeling failure.","tokens_in":28737,"tokens_out":3136,"duration_ms":29687,"concrete_test":"Hold out the pion data. Fix all simulation parameters using only electron and muon runs: the tile gap from the electron linearity fit, the SiPM saturation model from the electron energy-deposition distributions, and eps_scale for e- and mu- from their total-energy peaks. Then predict the 3 GeV pi- total-energy distribution and longitudinal profile using the standard Geant4 hadronic physics list, without applying the 0.82 pion scale factor or the 10%/35% contamination fractions. If the predicted pion peak is more than ~10% away from the data, the 0.82 correction is absorbing a real hadronic-modeling deficiency, and the 'validates the full simulation' claim should be weakened to 'the detector meets target performance when the simulation is tuned to test-beam data.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims the results 'validate the ENUBET full simulation.' In practice, the data-MC agreement used to support this claim is obtained only after several data-driven adjustments: a 1 mm inter-tile gap (Sec. 5.5), an average SiPM saturation model that the authors state overestimates saturation in the highest-deposit channel (Sec. 5.4 and Fig. 24), per-particle energy-scale corrections eps_scale = 0.97/0.82/1.0 for e-/pi-/mu- (Tab. 4 and Sec. 5.6), and fitted 10%/35% muon/pion contamination fractions in the hadron and muon samples (Sec. 5.6). The pion correction is especially large: the simulated pion energy is reduced by 18% after all other effects are included. These parameters are derived from the same total-energy distributions that are then shown to agree with data, so the comparison in Figs. 30-32 is a closure test rather than an independent validation. The residual layer-by-layer agreement is only 10-20%, with remaining discrepancies attributed to channel-dependent saturation variations that are not modeled. Because the full simulation is the basis for the quoted neutrino-flux systematics and PID performance, the abstract's 'validate' overstates the epistemic weight of this tuned comparison. The most load-bearing premise is that these parameters absorb benign, understood effects; if instead they are compensating for a genuine modeling failure (especially in hadronic response), the inferred validation and the 1% flux-systematics claim lack support from this test-beam campaign.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports the construction, commissioning, and test-beam performance of the ENUBET Demonstrator, a full-scale longitudinal sampling calorimeter prototype for the instrumented decay tunnel of a monitored neutrino beam. Using electron, pion, and muon beams at the CERN T9 line (0.5–10 GeV/c, with the main analysis at 1–5 GeV), the authors measure the electron energy resolution σ/E = (14.92 ± 0.21)%/√E ⊕ 1.08 MeV/E ⊕ (4.82 ± 0.25)%, a MIP light yield of 306–372 photoelectrons (Tab. 2), channel cross-talk below 5% (Fig. 21), and demonstrate e/π/μ separation through total-energy and longitudinal-shower-profile comparisons with a Geant4 simulation. They conclude that the results meet the requirements for neutrino monitoring in the ENUBET decay tunnel and that they validate the ENUBET full simulation.","tokens_in":29008,"tokens_out":5374,"duration_ms":44900,"significance":"If the results hold, the paper would provide strong experimental support for the scalability of the iron–scintillator–WLS–SiPM calorimeter technology to the ENUBET decay-tunnel instrumentation and would underwrite the 1% neutrino-flux systematic claim made by the collaboration. Strengths of the paper include the multi-year beam-test program, the internal consistency of the headline measurements (resolution, light yield, cross-talk), the careful channel-by-channel equalization, and the inclusion of a detailed detector geometry in the simulation. The main weakness is that the 'validation of the ENUBET full simulation' claim — which is central to the abstract and conclusions — rests on a simulation whose parameters are adjusted to the same data used for the comparison, making the agreement a closure test rather than an independent validation. This concern, together with the untested extrapolation to the high-rate tunnel environment, prevents me from recommending acceptance in the current form.","major_comments":[{"comment":"The abstract states that the results 'validate the ENUBET full simulation.' In Sec. 5.6 the MC agreement is obtained only after applying per-particle energy-scale corrections (ε_scale = 0.97/0.82/1.0 for e−/π−/μ−, Tab. 4), adding 10% muon contamination to the pion sample and 35% pion contamination to the muon sample, and incorporating a 1 mm tile gap (Sec. 5.5) and an average SiPM saturation model that the paper itself identifies as overestimating saturation in the highest-deposit channel (Sec. 5.4, Fig. 24). Because these parameters are derived from the same total-energy distributions that are then shown to agree, the comparison in Figs. 30–32 is a closure test, not an independent validation. The 'validate' language is therefore overstated. I request that the authors either soften the claim (e.g., 'the simulation describes the data after these adjustments') or provide an independent che","section":"Abstract, Sec. 5.6, Tab. 4, Figs. 30–32"},{"comment":"The MC energy-resolution prediction is obtained by averaging the MC results with and without SiPM saturation and assigning the spread as a systematic uncertainty. Since the saturation model is already known to overestimate the effect in the highest-energy channel (Sec. 5.4), this averaging does not constitute a validated prediction; it only brackets a known model deficiency. The authors should explicitly state that the MC resolution band is an estimate of model uncertainty, not a validated simulation result, and should adjust the wording of the abstract and conclusions accordingly. This is load-bearing because the 'full simulation validation' claim is one of the two central assertions of the paper.","section":"Sec. 5.5, Fig. 29"},{"comment":"The paper claims 'full scalability' of the technology and that the Demonstrator 'meets the requirements for neutrino monitoring' in the ENUBET decay tunnel, where rates of O(100–1000) kHz/cm² are expected. However, all beam-test data were taken with a low-rate, single-particle triggered beam (Sec. 4); no rate-dependent, pile-up, or simultaneous-multi-track measurement is presented. The extrapolation to the tunnel-rate environment is an untested assumption. Please explicitly state that the rate capability is not tested in this work and is deferred to future studies, or present a rate/pile-up measurement. As written, the claim of full scalability goes beyond the data shown.","section":"Sec. 1, Sec. 6"}],"minor_comments":[{"comment":"Typo: 'The total number of WLs fibers' should read 'WLS fibers.' Also, the sentence 'Former being constructed and tested in 2022, while the latter was constructed in 2023' should be 'The former was constructed and tested in 2022, while the latter was constructed in 2023.'","section":"Sec. 2, paragraph 2"},{"comment":"The text states that the relative uncertainty on the MIP MPV is approximately 10% for properly illuminated channels. Please specify what fraction of channels required the fallback 'highest bin' procedure and whether the quoted 10% includes the fallback cases.","section":"Sec. 5.2, Fig. 19"},{"comment":"The footnote about the π− run with a failed silicon tracker is useful but its impact on the fiducial selection and on the shower-profile comparison should be quantified or at least discussed more explicitly in the text.","section":"Sec. 5.6, footnote 1"},{"comment":"The right panel is described in the caption as 'Energy ratio between data and MC' but the text says it shows both e− and π−; the marker/legend for the two species is not clear in the figure. Please make the legend explicit.","section":"Fig. 32, right panel"},{"comment":"The phrase 'the electron component shows a pretty high purity' is informal for a journal paper; consider 'rather high purity' or 'high purity'.","section":"Sec. 5.6"},{"comment":"References [12] and [13] are marked 'in preparation'; if possible, update or add a note about the expected publication timeline, since the '1% flux systematic' claim in the conclusions relies on [12].","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid experimental detector paper, and the direct measurements (resolution, light yield, cross-talk) are internally consistent and well presented. The main barrier to acceptance is the 'validates the ENUBET full simulation' claim in the abstract and conclusions: the simulation is tuned to the data with several free parameters, so the agreement is a closure test. This is fixable by rewording and/or adding an independent validation, not by new data alone. The second concern — extrapolation to high-rate tunnel conditions — is also fixable by explicitly scoping the claim. I recommend major revision, not rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: the Demonstrator delivers the numbers ENUBET needs — electron energy resolution around 15%/√E ⊕ 4.8%, light yield 306–372 p.e./MIP, cross-talk below 5% — and those are direct, internally consistent measurements. The paper is worth reading if you care about monitored neutrino beams or finely segmented shashlik calorimeters. The four-fold light yield improvement over the 2020 prototype is concrete and the construction/commissioning detail is exactly what the field needs. Credit where it's due: the resolution, linearity, and MIP calibration are measured, not fitted to the conclusions, and the authors are unusually candid about imperfections (SiPM saturation overestimation, deferred fine-tuning, the broken tracker during the pion run, the 10–20% layer-by-layer discrepancies). That is honest reporting.\n\nThe soft spot is the abstract's claim that the data \"validate the ENUBET full simulation.\" The agreement is obtained only after several data-driven adjustments: a 1 mm tile gap, an average SiPM saturation model the authors admit overestimates saturation, per-particle energy-scale factors of 0.97/0.82/1.0, and fitted 10%/35% contamination fractions. The pion correction is large — an 18% reduction in simulated pion energy. Once you tune those parameters to the same total-energy distributions you then compare, the agreement in Figs. 30–32 is a closure test, not an independent validation. The stress-test note has this right. The layer-by-layer agreement at 10–20% also undercuts the word \"validate.\"\n\nTwo other things to keep in mind: nothing in this beam test reproduces the O(100–1000) kHz/cm² tunnel-rate environment that motivates the design, and the <1% flux systematic claim partly rests on an unpublished manuscript [12]. These are context issues, not fatal flaws. The detector-performance claims themselves don't depend on the simulation validation, so the core results stand.\n\nWho this is for: neutrino cross-section and oscillation people, calorimeter instrumentation folks, and anyone evaluating ENUBET/nuSCOPE. It deserves a serious referee. A good referee should push for a revised abstract that says \"supports\" or \"is consistent with\" rather than \"validates,\" plus a clearer accounting of how the tuned parameters were chosen and how robust the conclusions are to their variation. The underlying measurements should not be delayed for that.\n\nRecommendation: send it to peer review, with revision required on the validation claim.","headline":"Solid, honest detector paper whose direct performance numbers are trustworthy, but the abstract overstates the simulation validation — the comparison is a closure test with tuned parameters, not an independent check.","tokens_in":30090,"tokens_out":1704,"would_cite":true,"duration_ms":16694,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["29.40.Vj","29.40.Mc"],"model":"deepseek-v4-flash","headline":"A full-size prototype of the ENUBET instrumented decay tunnel meets the performance needed to monitor neutrino beams, validating the detector technology and its simulation.","keywords":["calorimeters","neutrino detectors","monitored neutrino beams","particle identification","silicon photomultipliers","energy resolution","test beam","scintillator tiles"],"falsifier":"A concrete test would be to build a small module with precisely machined tiles (known coating thickness, no manual glue variations) and no free parameters: measure the MIP response, electron shower profiles, and energy resolution, and compare directly with the unadjusted Monte Carlo. If the simulation then fails without the 1 mm gap or the 0.82 pion scale factor, the validation claim would be weakened. Alternatively, operating the Demonstrator under an intense beam that reproduces the O(100–1000) kHz/cm² rates would test the high-rate extrapolation.","tokens_in":28485,"feed_emoji":"⚛️","tokens_out":4753,"duration_ms":38090,"temperature":0.7,"pith_summary":"The paper reports on the construction and test-beam performance of a full-size prototype of the instrumented decay tunnel for a 'monitored neutrino beam,' a design that measures the neutrinos' parents (electrons, muons, pions) as they decay in the tunnel. The authors show that a longitudinally segmented iron-and-scintillator calorimeter, read out with wavelength-shifting fibers and silicon photomultipliers, achieves an electron energy resolution of about 15%/√E plus a small constant term, comfortably inside the 25%/√E requirement. They also demonstrate electron/pion/muon separation using total-energy and shower-profile information, and report a fourfold improvement in light yield over an earlier prototype. A detailed simulation reproduces the recorded distributions once a few physically motivated adjustments are included, which the authors take as validation of the full ENUBET simulation. A sympathetic reader would care because this is the experimental evidence that monitored neutrino beams can reach the sub-1% flux systematic uncertainty demanded by next-generation neutrino cross-section experiments.","feed_headline":"Full-size calorimeter prototype passes neutrino-beam monitor test","feed_subtitle":"Test-beam data show electron/pion separation and energy resolution good enough for 1% neutrino-flux uncertainty.","key_machinery":"The key object is the 'Lateral-readout Compact Module' (LCM): a calorimeter cell made of five iron slabs interleaved with plastic scintillator tiles (~3×3 cm²), each tile read by two wavelength-shifting fibers glued in grooves, with fibers from a radial column bundled into a single silicon photomultiplier after passing through borated-polyethylene neutron shielding. Longitudinal sampling every 4.3 radiation lengths, together with transverse granularity, provides the total-energy and shower-shape information used to separate electrons, pions, and muons. The simulation machinery is a full Monte Carlo detector description that includes tile geometry, fiber grooves, and a configurable tile gap,","core_discovery":"The central claim is that the Demonstrator — a 1.7 m-long slice of the ENUBET decay-tunnel instrumentation — meets the performance requirements for lepton identification, and that the detector's full simulation is validated by the test-beam data. Concretely: electron energy resolution σ/E = (14.92±0.21)%/√E ⊕ 1.08 MeV/E ⊕ (4.82±0.25)% meets the <25%/√E design value; MIP light yield is 306–372 photoelectrons, about four times the 2020 prototype; channel cross-talk is below 5%; and the total-energy and longitudinal-shower profiles for e−, π−, μ− are reproduced by the simulation within 10–20% once a 1 mm tile gap, an average SiPM saturation model, per-particle energy-scale factors, and small co","pith_inferences":["The reliance on per-particle energy-scale factors (0.97, 0.82, 1.0) and fitted contamination fractions suggests the simulation's agreement with data is partly absorbed by free parameters; a stricter test would use a prototype with independently measured tile gaps and optical coupling to see if the same adjustments remain necessary.","The SiPM saturation observed at a few GeV is an early warning for the full detector: at the higher rate environment of a real tunnel, saturation may be worse, motivating smaller SiPM cells or a different readout strategy.","The test-beam data is single-particle and low-rate; extending to the O(100–1000) kHz/cm² tunnel environment will require a dedicated high-rate test to confirm the pile-up and timing performance.","If the simulation validation holds, the same modeling approach can be applied to other proposed monitored-beam detectors (e.g., neutrino cross-section experiments at similar facilities)."],"forward_implications":["If correct, the same modular technology can be scaled to instrument the full decay tunnel at the stated cost, enabling monitored neutrino beams with flux systematic uncertainties below 1%.","The demonstrated e/π/μ separation means the detector can count positrons from K→eνπ⁰ and muons from π/K→μν, providing an event-by-event neutrino flux monitor.","The fourfold light yield improvement shows the revised optical-coupling scheme (frontal tile readout, better fiber gluing) works, and points the way to further gains.","The simulation, once adjusted, can be used to predict the performance of the final detector and to design the readout electronics for the high-rate tunnel environment."],"fun_headline_variants":["ENUBET demonstrator meets lepton ID specs in beam tests","Full-size tunnel calorimeter passes neutrino monitor test","Test-beam data validate ENUBET full-size calorimeter","Electron resolution meets design in ENUBET prototype","ENUBET demonstrator: beam tests confirm simulation"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The claim that the full simulation is 'validated' rests on the assumption that the data-driven adjustments—the 1 mm tile gap, the average SiPM saturation model, the per-particle energy-scale factors, and the fitted contamination fractions—represent genuine, understood physical effects rather than flexible parameters that absorb real modeling failures; a further assumption is that low-rate single-particle test beams reproduce the high-rate tunnel environment.","fun_headline_variants_meta":{"raw":{"variants":["ENUBET demonstrator meets lepton ID specs in beam tests","Full-size tunnel calorimeter passes neutrino monitor test","Test-beam data validate ENUBET full-size calorimeter","Electron resolution meets design in ENUBET prototype","ENUBET demonstrator: beam tests confirm simulation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000177,"raw_usage":{"total_tokens":1122,"prompt_tokens":728,"completion_tokens":394,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":472,"completion_tokens_details":{"reasoning_tokens":315}},"tokens_in":472,"tokens_out":394,"duration_ms":4856,"temperature":1.0,"reasoning_tokens":315,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T13:28:08.212900+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test would be to build a small module with precisely machined tiles (known coating thickness, no manual glue variations) and no free parameters: measure the MIP response, electron shower profiles, and energy resolution, and compare directly with the unadjusted Monte Carlo. If the simulation then fails without the 1 mm gap or the 0.82 pion scale factor, the validation claim would be weakened. Alternatively, operating the Demonstrator under an intense beam that reproduces the O(100–1000) kHz/cm² rates would test the high-rate extrapolation.","supporting_citations":[],"review_version":1}