{"id":"54fa9d0f-7073-4e6f-bbdd-dcb9b637c503","arxiv_id":"2411.09562","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"In simulated IWCD data, a ResNet-18 classifier selects electron neutrino events with 61.5% purity and 78.2% efficiency, improving on fiTQun's 51.1% and 69.5%.","lead":"This paper compares two ways to pick out electron neutrino events in the planned Intermediate Water Cherenkov Detector (IWCD) for Hyper-Kamiokande: a standard likelihood algorithm and a deep learning classifier. In simulated data, the deep learning method achieves higher purity and efficiency for the electron neutrino sample.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported ML purity/efficiency are in-sample estimates because the three discriminators were manually tuned on the same beam sample used for Table 1; the comparison with fiTQun may be optimistic.","rationale":"The reader's verdict identified the same load-bearing weakness: the ML cuts are manually tuned on the same simulated beam sample used to report the final purity and efficiency, making the improvement an in-sample estimate with no demonstration of generalization. This is indeed the most consequential concern because the paper's entire quantitative conclusion is the Table 1 comparison. My independent check of the table's arithmetic confirms internal consistency (the 51.1% and 61.5% purities follow from the listed counts if 'Total NC' is a superset of NCπ0 and NCγ), so I do not suspect a numerical error. The absence of uncertainties and the lack of released code or data are secondary but compound the concern: without a validation split, there is no way to quantify how much of the reported gain is due to selection tuning. The proposed held-out test is the minimal, decisive check. If the authors can show that the held-out purity/efficiency remain above the fiTQun values, the central claim would be materially strengthened. Until then, the conditional verdict is appropriate: the result is plausible but not yet established. I therefore recommend keeping the reader's CONDITIONAL verdict unchanged.","tokens_in":5164,"tokens_out":4390,"duration_ms":39851,"concrete_test":"Before any further tuning, randomly split the simulated beam sample into a tuning set (80%) and a held-out validation set (20%). Re-tune the three ML discriminators on the tuning set only (e.g., by maximizing FOM on that subset), freeze the cuts, then apply them to the held-out set to compute purity and efficiency. If the held-out values are consistent with the in-sample numbers within statistical uncertainty (e.g., purity ≥ 58% and efficiency ≥ 74%), the claim survives. If they fall to or below the fiTQun values (51.1% purity, 69.5% efficiency) or drop substantially below 61.5%/78.2%, the reported ML improvement is largely an artifact of tuning on the evaluation sample.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that a ResNet-18 based selection achieves 61.5% purity and 78.2% efficiency versus fiTQun's 51.1% and 69.5%—rests on numbers computed on the same simulated beam sample used to tune the ML selection cuts. Section 3 states: 'Based on the distribution of signal and background in the histograms, we manually tuned three discriminators across P(μ), P(π0), and P(e) to improve both the purity and efficiency.' Those histograms (Figure 4) are produced from the 'test dataset' used for ML model evaluation, and Table 1's final row is quoted from that same sample after applying the manually tuned cuts. No independent validation event sample, cross-validation, or uncertainty estimate is mentioned. Because thresholds are chosen to maximize purity/efficiency on the evaluation set, the reported 61.5%/78.2% are in-sample estimates and are expected to degrade on a fresh sample. The fiTQun cuts were also optimized on the same sample, but with a predefined Figure-of-Merit (FOM = S/√(S+B)), so the two methods are not compared on equal footing: the ML result includes an extra hand-tuning step that can absorb statistical fluctuations in the simulated dataset. This is the most load-bearing concern because if an unbiased evaluation shrinks or reverses the gap, the paper's main physics message disappears. The internal arithmetic of Table 1 is consistent (e.g., fiTQun: 2535+144+2100+24+157+1 = 4961, 2535/4961 = 51.1%), so the issue is not computation but statistical validity of the comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This ICHEP 2024 proceedings paper compares two event-selection methods for electron-neutrino charged-current quasi-elastic (nu_e CC0pi) events in the Intermediate Water Cherenkov Detector (IWCD) of Hyper-Kamiokande, using simulated data from NEUT, WCSim, and fiTQun. The fiTQun-based analysis applies fiducial-volume, kinematic, and likelihood-ratio cuts and reports a purity of 51.1% and efficiency of 69.5%. The authors then train a ResNet-18 CNN within the WatChMaL framework on particle-gun simulated events and apply it to the same simulated beam sample; after manually tuning three softmax-probability discriminators, they report an improved purity of 61.5% and efficiency of 78.2%. The paper concludes that the ML-based selection outperforms fiTQun and expects further gains with automated cut optimization.","tokens_in":5535,"tokens_out":8966,"duration_ms":78739,"significance":"If the reported improvement is robust, it is a useful result for Hyper-Kamiokande's IWCD physics program: better nu_e event selection would reduce backgrounds for cross-section measurements and, ultimately, for CP-violation sensitivity. The paper provides concrete simulated event counts and a direct head-to-head comparison with an established likelihood-based reconstruction, which is valuable to the community. The main weakness is statistical: the ML cut thresholds are manually tuned on the same simulated beam sample used to quote the final purity and efficiency, so the headline numbers are in-sample estimates. A proper validation on a held-out sample, or cross-validation, is needed before the central claim can be accepted as stated.","major_comments":[{"comment":"The central claim that the ResNet-18 selection achieves 61.5% purity and 78.2% efficiency is based on in-sample evaluation. The paper states that the three discriminators were manually tuned 'based on the distribution of signal and background in the histograms' (Figure 4), and that the same 'test dataset' is used for ML model evaluation. The final purity and efficiency in Table 1 are then quoted from that same sample after applying the tuned cuts. No independent validation split, cross-validation, or hold-out set is described. This is a load-bearing issue because manual threshold tuning can exploit statistical fluctuations in the simulated sample, making the reported improvement over fiTQun optimistic. Please evaluate the selected sample on events that were not used for tuning the ML cuts, and report the purity/efficiency with statistical uncertainties.","section":"Section 3, Figure 4, Table 1"},{"comment":"The comparison between fiTQun and ML is not made on equal footing. The fiTQun cut lines are optimized using a predefined figure of merit, FOM = S/sqrt(S+B), whereas the ML softmax cuts are manually tuned to improve both purity and efficiency without an explicit, common objective. This asymmetry means that the ML result includes an additional hand-tuning step that can absorb favorable statistical fluctuations. A fairer comparison would either optimize both methods with the same FOM on a training sample and evaluate both on a separate test sample, or use a blinded analysis strategy.","section":"Sections 2.1 and 3, Table 1"},{"comment":"The central table is not self-explanatory as printed. The header appears to list more categories than can be unambiguously matched to the numeric entries, and the sum of the listed counts does not directly reproduce the denominators implied by the quoted purities (e.g., for the ML row, the listed categories sum to 5362 events, while the quoted purity of 61.5% corresponds to a denominator of about 4642 events). Please restate Table 1 with explicit column definitions, a clear 'total selected' row/column, and the exact formula used for the purity denominator, so that the results are reproducible.","section":"Table 1"}],"minor_comments":[{"comment":"The terminology 'test dataset' is confusing: the same sample is used both for manually tuning the discriminators and for reporting the final numbers. Please use consistent terms such as training/validation/test, or describe the split explicitly.","section":"Section 3"},{"comment":"No statistical or systematic uncertainties are reported for the event counts, purity, or efficiency. At minimum, Poisson or binomial errors on the MC counts should be given, and the absence of detector/systematic uncertainties should be stated as a limitation.","section":"Table 1"},{"comment":"The input representation for the ResNet-18 model is not described: what exactly is fed to the network (mPMT hit images, charge/time channels, etc.)? Also, details of the train/validation split used during the 20-epoch training are omitted.","section":"Section 3"},{"comment":"The simulation chain is not fully referenced: NEUT and WCSim are mentioned in the text but no citations are provided. Please add the appropriate references.","section":"Section 2"},{"comment":"The statement that fiTQun processes 'at most 1 event per minute' is a quantitative runtime claim without supporting measurements or hardware specifications. If computational cost is part of the motivation, provide a benchmark on the same hardware for both methods.","section":"Section 2.1"},{"comment":"Some figure references and caption details are incomplete in the extracted text (e.g., axis labels, color scales, and exact definitions of 'Total NC' versus 'NC pi0' and 'NC gamma'). Please ensure the final version includes clear, self-contained figure and table captions.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This is a conference proceedings with a plausible and potentially useful result, but the central comparison is weakened by in-sample tuning and the table is not fully reproducible as printed. The authors should be asked to add a validation set or cross-validation, clarify Table 1, and include uncertainties. If these are addressed, the paper could be acceptable as a proceedings contribution; the present version is not yet quantitatively reliable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a short ICHEP proceeding reporting that a ResNet-18 CNN trained in the WatChMaL framework selects electron neutrino CC0π events at IWCD with 61.5% purity and 78.2% efficiency, versus fiTQun's 51.1% and 69.5% on the same simulated beam sample. The new result is the specific quantitative comparison for IWCD; the method itself is an extension of established work, and the authors cite it properly. The internal arithmetic in Table 1 checks out, the signal definition is clear, and the particle-gun training set for the CNN is independent of the beam sample, which gives the network some grounding.\n\nThe soft spot is exactly where you'd expect for a conference short paper: the three softmax cut thresholds (P(e), P(μ), P(π0)) were manually tuned on the same simulated beam dataset used to quote the final Table 1 numbers. No validation split, no cross-validation, no uncertainties. That makes the reported improvement an in-sample estimate, probably optimistic. fiTQun's cuts were also optimized on the same sample, so both are in-sample, but the ML manual tuning is more flexible and can absorb statistical fluctuations. The authors do not report the cut values, nor do they release code or data, so the comparison is not independently reproducible. The AUC comparison in Figure 3 (0.7183 vs 0.5418 for gamma rejection) is suggestive, but it suffers from the same evaluation-set issue.\n\nNone of this makes the paper worthless. As a proceedings contribution it is a reasonable status report from the Hyper-K collaboration, and the direction is sound. But a reader should not take the 61.5%/78.2% numbers as established performance. The fix is straightforward: hold out a validation sample (or use k-fold) and report the cuts and their uncertainties. If the improvement survives that, it's a useful result for cross-section and CPV studies; if it shrinks, the paper's main message changes.\n\nI would send it to a serious referee rather than desk reject, because the topic matters and the flaw is fixable. But I would not cite the specific numbers in my own work until they are validated on a fresh sample. It's a useful reading-group example of in-sample tuning in ML-based HEP analyses.","headline":"A plausible but in-sample ML improvement for IWCD electron neutrino selection; the reported purity/efficiency gains are likely optimistic until validated on a fresh sample.","tokens_in":6105,"tokens_out":3949,"would_cite":false,"duration_ms":35268,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A deep-learning event selection for the Intermediate Water Cherenkov Detector raises simulated electron-neutrino sample purity from 51.1% to 61.5% while improving efficiency from 69.5% to 78.2% compared to the likelihood-based fiTQun…","keywords":["electron neutrino CC0pi selection","Intermediate Water Cherenkov Detector","Hyper-Kamiokande","convolutional neural network","particle identification","fiTQun","NC pi0 background","purity and efficiency"],"falsifier":"Train the identical network and apply the identical softmax cut thresholds to a freshly generated independent IWCD Monte Carlo sample, then recompute purity and efficiency; if the numbers fall materially below 61.5% and 78.2%, the reported gain is overfit to the tuning sample. A simpler check is to split the existing simulated beam sample in half, tune on one half and evaluate on the other.","tokens_in":4990,"feed_emoji":"⚛️","tokens_out":10338,"duration_ms":83658,"temperature":0.7,"pith_summary":"The paper aims to show that a modern convolutional neural network can select electron-neutrino charged-current events without pions more cleanly than the standard likelihood-based reconstruction used in water Cherenkov detectors. For the Intermediate Water Cherenkov Detector (IWCD) planned for the Hyper-Kamiokande program, the authors report that their machine-learning selection raises the purity of the simulated $\\nu_e$CC0$\\pi$ sample from 51.1% to 61.5% and the efficiency from 69.5% to 78.2% compared with the fiTQun likelihood algorithm. These gains matter because electron neutrinos make up only about 1.5% of the accelerator beam, so better rejection of backgrounds such as neutral-current $\\pi^0$ production directly reduces the systematic uncertainties that limit measurements of CP violation in the neutrino sector. The paper presents the improvement as evidence that machine learning can outperform likelihood-based particle identification for this detector.","feed_headline":"Deep learning lifts electron-neutrino purity to 61.5%","feed_subtitle":"On simulated IWCD data, ResNet-18 lifts purity and efficiency, 61.5% and 78.2% vs 51.1% and 69.5%.","key_machinery":"The mechanism that carries the argument is the trained softmax output of an 18-layer residual convolutional network (ResNet-18). The network converts raw PMT hit patterns into four class probabilities $P(e)$, $P(\\mu)$, $P(\\gamma)$, and $P(\\pi^0)$; the selection then applies manually chosen two-dimensional cuts on these probabilities versus reconstructed momentum, layered on the same fiducial-volume and kinematic cuts used in the fiTQun analysis. The baseline against which the claim is measured is fiTQun, a maximum-likelihood event reconstruction whose likelihood ratios are used for particle identification.","core_discovery":"On the paper's own terms, the central discovery is that a ResNet-18 convolutional network trained on four particle classes ($e^-$, $\\mu^-$, $\\gamma$, $\\pi^0$) yields softmax probabilities that, after manual cuts tuned against reconstructed momentum, select $\\nu_e$CC0$\\pi$ events at higher purity and efficiency than fiTQun's likelihood ratios on the same simulated beam sample. The ML selection retained 2856 signal events versus 2535 for fiTQun, achieved 61.5% purity and 78.2% efficiency, and suppressed NC$\\pi^0$ and NC$\\gamma$ backgrounds more strongly, at the cost of admitting more $\\nu_e$CC-other and $\\bar{\\nu}_e$ events. The authors conclude that with further development, more complex and automated machine-learning cuts should improve the sample beyond what is reported.","pith_inferences":["The reported numbers are in-sample estimates because the softmax cuts were manually tuned on the simulated beam sample used to evaluate them; an independent test sample would likely show a smaller (though possibly still positive) improvement.","The network was trained only on single-particle particle-gun events, not on full neutrino interaction topologies; retraining or fine-tuning on beam-like events with multiple rings could improve rejection of asymmetric $\\pi^0$ decays further.","The same softmax-plus-cut strategy could transfer to Hyper-K's far detector, but the different PMT geometry and granularity make that transfer an open question.","Because the ML selection increases the $\\nu_e$CC-other and $\\bar{\\nu}_e$ backgrounds, the net benefit for a cross-section measurement depends on how well those backgrounds are constrained independently."],"forward_implications":["A convolutional-network PID can serve as the $\\nu_e$ event selection for IWCD, replacing likelihood-based cuts with a single pass through a trained classifier.","The higher purity and efficiency reduce the statistical and systematic footprint of the $\\nu_e$ cross-section measurement, which is one of the inputs to the CP-violation sensitivity of the Hyper-Kamiokande program.","Because NC$\\pi^0$ and NC$\\gamma$ contamination is the hardest part of the likelihood selection, the ML result shows these backgrounds can be suppressed substantially without losing signal.","The authors expect further development of more complex and automated machine-learning cuts to push the sample's purity and efficiency beyond the reported values."],"supporting_citations":[{"why":"Describes the Intermediate Water Cherenkov Detector whose simulated event sample is analyzed here.","marker":"[5]"},{"why":"Defines the original likelihood-based fiTQun reconstruction and particle identification algorithm used as the baseline.","marker":"[6]"},{"why":"Derives the maximum-likelihood formulation for water Cherenkov reconstruction that fiTQun relies on.","marker":"[7]"},{"why":"Demonstrates the improved event reconstruction in Super-Kamiokande that motivates the likelihood approach.","marker":"[8]"},{"why":"Introduces the ResNet-18 residual convolutional network architecture used for the machine-learning classifier.","marker":"[11]"},{"why":"Provides the machine-learning framework used to train and evaluate the convolutional network on detector images.","marker":"[12]"},{"why":"Earlier application of machine learning to water Cherenkov event reconstruction that motivates the present analysis.","marker":"[13]"}],"fun_headline_variants":["Deep learning boosts electron-neutrino selection in Hyper-K's IWCD","ResNet-18 lifts electron neutrino efficiency to 78.2% in Hyper-K IWCD","Machine learning beats likelihood for electron neutrino events at IWCD","Hyper-K IWCD: deep learning selected nu_e events with 61.5% purity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The cuts producing the 61.5% purity and 78.2% efficiency were tuned by eye on the same simulated sample from which the final numbers are drawn, so the result assumes that sample is a faithful, unbiased stand-in for real IWCD data and that the tuning did not overfit its statistical fluctuations.","fun_headline_variants_meta":{"raw":{"variants":["Deep learning boosts electron-neutrino selection in Hyper-K's IWCD","ResNet-18 lifts electron neutrino efficiency to 78.2% in Hyper-K IWCD","Machine learning beats likelihood for electron neutrino events at IWCD","Hyper-K IWCD: deep learning selected nu_e events with 61.5% purity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000906,"raw_usage":{"total_tokens":3901,"prompt_tokens":952,"completion_tokens":2949,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":568,"completion_tokens_details":{"reasoning_tokens":2860}},"tokens_in":568,"tokens_out":2949,"duration_ms":21383,"temperature":1.0,"reasoning_tokens":2860,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:30:43.554523+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the identical network and apply the identical softmax cut thresholds to a freshly generated independent IWCD Monte Carlo sample, then recompute purity and efficiency; if the numbers fall materially below 61.5% and 78.2%, the reported gain is overfit to the tuning sample. A simpler check is to split the existing simulated beam sample in half, tune on one half and evaluate on the other.","supporting_citations":[{"cited_title":"An intermediate water cherenkov detector at j-parc","cited_arxiv_id":null,"evidence_quote":"Describes the Intermediate Water Cherenkov Detector whose simulated event sample is analyzed here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the original likelihood-based fiTQun reconstruction and particle identification algorithm used as the baseline."},{"cited_title":"Improving the t2k oscillation analysis with fitqun: a new maximum-likelihood event reconstruction for super-kamiokande","cited_arxiv_id":null,"evidence_quote":"Derives the maximum-likelihood formulation for water Cherenkov reconstruction that fiTQun relies on."},{"cited_title":"Atmosphericneutrinooscillationanalysiswithimprovedeventreconstruction in super-kamiokande iv.Progress of Theoretical and Experimental Physics, 2019(5):053F01, 2019","cited_arxiv_id":null,"evidence_quote":"Demonstrates the improved event reconstruction in Super-Kamiokande that motivates the likelihood approach."},{"cited_title":"Watchmal: Github repository.https://github.com/WatChMaL/WatChMaL","cited_arxiv_id":null,"evidence_quote":"Provides the machine-learning framework used to train and evaluate the convolutional network on detector images."},{"cited_title":"Machine learning techniques to enhance event reconstruction in water cherenkov detectors","cited_arxiv_id":null,"evidence_quote":"Earlier application of machine learning to water Cherenkov event reconstruction that motivates the present analysis."}],"review_version":1}