{"id":"38827bb9-62cd-4032-b915-c4f6952d70a7","arxiv_id":"2501.08671","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A set of analysis upgrades for monoscopic IACT events improves angular resolution by 57% at 100 GeV, lowers the energy threshold by about half, and increases low-energy sensitivity by 41% in simulations.","lead":"This paper improves how single-telescope gamma-ray events are analyzed at the H.E.S.S. observatory, using machine learning to determine shower orientation, new image parameters, and an intensity-dependent selection cut. If the improvements hold, they could lower the energy threshold and boost sensitivity for low-energy gamma-ray astronomy, which is relevant for future observatories like CTAO.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported gains may be inflated because the size-dependent cut, pixel-number preselection, and variable selection are optimized on the same simulated sample used for the headline performance numbers, with no fully independent test set documented.","rationale":"I read the paper as a methods paper whose central claim is quantitative: the new chain improves low-energy monoscopic performance by the stated margins. The claim stands or falls on whether those margins are real. The paper does several things well: it shows variable distributions against real background data (Appendix A), estimates NSB systematics with a 1.65x NSB simulation, and applies the analysis to Crab data with consistent results. The Crab application is genuine validation, but it is a consistency check above 100 GeV, not a direct measurement of the 30 GeV threshold or the 41% sensitivity gain. The main risk is therefore internal: the same simulated events are used both to optimize the analysis (so the analysis is partly selected to look good on those events) and to report performance. This is a well-known overfitting route. The reader's weakest assumption, simulation fidelity, is real but is a physics assumption that cannot be fully settled within this paper; the double-use issue is a methodological assumption that is directly checkable and should be the condition for acceptance. I therefore keep the verdict CONDITIONAL: the authors should document and, if necessary, re-derive the headline numbers with a strictly independent test set for the cut and preselection choices. The paper's code and data are not provided, which compounds the difficulty, but the requested split is a concrete, feasible check. Agreement with the reader is partial because the reader named simulation fidelity as the weakest assumption while also noting the size-dependent-cut issue in the rationale; my concern is the same internal validation issue, sharpened and elevated to the most load-bearing point.","tokens_in":24539,"tokens_out":6641,"duration_ms":65121,"concrete_test":"Split the simulated gamma-ray events and the measured background events into two independent halves before any optimization. Use half A for: (i) flip/energy/gammaness training and iterative variable selection, (ii) the q-factor scan in Sect. 2.3 to derive the smoothed size-dependent BDT cut, and (iii) the pixel-number preselection scan in Sect. 3.3. Then compute the effective-area threshold, angular-resolution ratio, and sensitivity ratio of Sect. 3 entirely on half B, with no further parameter changes. Compare to the published values; if the 41% low-energy sensitivity gain drops below ~25% or the 10%-effective-area threshold rises above ~45 GeV, the published improvements are substantially overfit.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline numbers (57% angular-resolution gain at 100 GeV, 41% sensitivity gain at threshold, 30 GeV effective-area threshold) are all computed on the same simulated sample that is used to choose the analysis configuration. Section 2.3 optimizes the size-dependent BDT cut by scanning q = eps_sig/sqrt(eps_bkg) in size bins over simulated events. Section 3.3 selects the 'seven or more pixels' preselection by scanning sensitivity over simulations. Section 2.2 performs iterative variable selection using validation losses/efficiencies. Although the ML models themselves use held-out validation (10% for separation, 20% for flip), the q-factor scan and the pixel-number cut are not documented as using a separate, untouched sample. Because the same events determine both the optimal cuts and the reported performance, the optima are subject to selection bias: statistical fluctuations in the low-size bins can be exploited by the cut scan, inflating the effective area at threshold and the sensitivity gain. The 41% and 30 GeV figures are therefore not protected against overfitting by the procedure described in the paper. This is a concrete, internal methodological risk that can be checked, unlike the broader simulation-fidelity assumption.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents three improvements to monoscopic IACT analysis using H.E.S.S. CT5/FlashCam data: a BDT-based determination of the image orientation (flip), roughly 30 new image parameters characterizing intensity and time distributions, and a size-dependent gamma-hadron selection cut. The authors report a 57% improvement in angular resolution at 100 GeV, a reduction of the 10% effective-area energy threshold from 64 GeV to 30 GeV, and a 41% improvement in sensitivity at the low-energy threshold. The methods are evaluated on CORSIKA/sim_telarray simulations and validated on ~30 h of Crab Nebula data.","tokens_in":24799,"tokens_out":4906,"duration_ms":47483,"significance":"If the reported gains are robust, this is a valuable contribution to monoscopic IACT analysis, with direct relevance to H.E.S.S. and to the design of monoscopic analyses for CTAO. The paper is careful in several respects: the ML models use held-out validation sets, the NSB sensitivity of each variable is quantified with an Anderson-Darling test, the systematic impact of a 1.65x higher NSB is evaluated for key IRFs, and the new analysis is applied to real Crab data. These strengths make the central idea credible. However, the quantitative headline claims are simulation-based, and the manuscript does not document a fully independent evaluation of the analysis configuration, which introduces a selection-bias risk into the reported improvements.","major_comments":[{"comment":"The size-dependent BDT cut (Section 2.3) and the 'seven or more pixels' preselection (Section 3.3) are optimized by scanning the same simulated sample that is later used to compute the effective areas and differential sensitivities in Figures 12-14. The paper does not state that a separate, untouched sample was reserved for the final performance evaluation. Because the q-factor scan selects the maximum over many size bins, statistical fluctuations in low-statistics bins can be exploited, potentially inflating the low-size effective area, the 30 GeV threshold, and the 41% sensitivity gain. Please either reserve a final evaluation sample that is not used for any cut/preselection choice, or demonstrate via cross-validation or a repeated split that the headline numbers are stable. Reporting the number of simulated events per size bin in Figure 6 would also help the reader assess the scale of fluctuations.","section":"Sections 2.3 and 3.3"},{"comment":"The iterative variable-selection procedures for the flip BDT, the regression NNs, and the separation BDT use validation losses/efficiencies from the same simulation set that produces the reported reconstruction and separation performance. While 10-20% of events are held out from training, the final variable subsets are selected on those validation sets, and the same overall sample then contributes to the headline numbers; this is a form of selection bias. The 57% angular-resolution improvement and the 20% q-factor improvement are the quantities most at risk. Please report the final configuration's performance on a sample that was not used for variable selection, hyperparameter choice, or early stopping, or quantify the selection variance by repeating the selection procedure on bootstrap subsamples.","section":"Sections 2.2.1-2.2.3"}],"minor_comments":[{"comment":"The sentence 'A low-energy threshold of 100 GeV is chosen, which lies above the 10% effective area threshold and below the 10% energy bias threshold for this zenith range' appears inconsistent with the immediately following statement that the new analysis exhibits an energy bias of 41% at 100 GeV; if the bias already exceeds 10% at 100 GeV, then 100 GeV lies above, not below, the 10% bias threshold.","section":"Section 3.4"},{"comment":"The y-axis label contains a typo: 'nromalized' should be 'normalized'.","section":"Figure 6"},{"comment":"The notation in the 'Mirror' column (e.g., '×(−1)', '−xhead', 'ytail yhead') is cryptic; a sentence in the table caption explaining the convention would improve readability.","section":"Table B.1"},{"comment":"The Anderson-Darling statistic for distance-top-2 is listed as -0.408; since this test statistic is usually nonnegative, please clarify how a negative value arises or correct the entry.","section":"Section 2.1.5 / Table B.1"},{"comment":"After the Gaussian smoothing of the per-bin optimal cut, the text says events are selected using the 'BDT cut interpolated at their respective size,' but it does not specify whether the interpolation is linear in log-size or how boundary bins are handled; please state the interpolation scheme explicitly.","section":"Section 2.3"},{"comment":"The data availability statement only provides the appendix figures on Zenodo; sharing the trained models, the optimized cut curve, and the analysis configuration would substantially improve reproducibility.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is well written and the methodological idea is sound, but the strongest quantitative claims are not yet protected against selection bias because the same simulation set is used both to choose the analysis configuration and to compute the headline performance numbers. I would support acceptance after the authors either document a fully independent evaluation sample or demonstrate with cross-validation that the reported gains are stable. The scope fits A&A and the paper is likely to be a useful reference for monoscopic IACT analyses."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a careful read. The combination of a BDT-based flip determination, about thirty new image parameters, and a size-dependent gamma-hadron cut gives large gains in simulation: 57% better angular resolution at 100 GeV, the effective-area threshold down from 64 to 30 GeV, and 41% better sensitivity at threshold. These are real, potentially important improvements for monoscopic analysis, which is the bulk of CT5 data below 100 GeV.\n\nWhat the paper does well: the flip BDT is a clean idea, and the head-tail sensitive variables make physical sense. The iterative variable selection is systematic, and the NSB sensitivity checks are more careful than most comparable papers. The Crab application is a useful sanity check, and the authors are clear that their baseline is not identical to Murach et al. (2015). They also hold out test sets for the ML models, which is more than many analysis papers do.\n\nThe soft spot is the one your stress-test note flags, and I think it is real. The size-dependent cut (Sec. 2.3), the 7-pixel preselection (Sec. 3.3), and the variable selection (Sec. 2.2) are all chosen using the same simulated sample that produces the headline numbers. The q-factor scan and the pixel-number scan are not documented as using an untouched test set. That means statistical fluctuations in the low-size bins, where Monte Carlo statistics are thinnest, can be exploited, inflating the 41% sensitivity gain and the 30 GeV threshold. The ML models themselves use validation splits, but the selection process peeks at those splits repeatedly, which is a milder version of the same bias. This is a concrete, checkable flaw, not a vague worry: a referee should ask for the cut and preselection to be validated on an independent simulation set, or for a nested cross-validation. The lack of released code or data makes this harder to audit.\n\nA second, lesser point: part of the threshold reduction comes from switching from a constant cut to a size-dependent cut, which by design retains more low-size events. The decomposition in Fig. 11 helps, but the headline numbers mix the cut change with the new variables.\n\nOverall, the central idea is sound and the paper deserves serious refereeing. The circularity concern is addressable and likely does not break the main conclusions, but it does mean the exact numbers should be treated with caution until confirmed on an independent sample. I would send this to peer review, with a request for a clean validation of the tuned cuts.","headline":"Solid, genuinely useful improvements to mono IACT reconstruction, but the headline gains may be partly inflated because the cut and preselection are optimized on the same simulations used to measure them.","tokens_in":25329,"tokens_out":3398,"would_cite":true,"duration_ms":35304,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A machine-learning overhaul of single-telescope gamma-ray analysis halves the low-energy threshold and sharpens angular resolution by 57 percent at 100 GeV.","keywords":["imaging atmospheric Cherenkov telescopes","monoscopic event reconstruction","gamma-hadron separation","Hillas parameters","boosted decision trees","H.E.S.S.","size-dependent selection cut","energy threshold"],"falsifier":"Measure the monoscopic Crab Nebula spectrum between 50 and 100 GeV with the new chain and compare it against the stereoscopic measurement of the same source: the paper claims an energy bias below 10% down to 50 GeV, so a mono spectrum departing from the stereo spectrum by more than the quoted systematic band in that range would show the simulated low-energy instrument response is wrong.","tokens_in":24353,"feed_emoji":"🔭","tokens_out":10735,"duration_ms":92683,"temperature":0.7,"pith_summary":"This paper claims that the weakest link in gamma-ray astronomy with imaging atmospheric Cherenkov telescopes—events recorded by a single telescope, which dominate at low energies—can be substantially strengthened by three coordinated changes to the analysis chain. It replaces the skewness-based decision about which end of an image points at the source (the 'flip') with a boosted decision tree, adds roughly thirty new image parameters that capture the shape, brightness structure, and time structure of the shower image beyond the classic Hillas moments, and makes the gamma-hadron selection cut depend on the total image intensity rather than being constant. On H.E.S.S. simulations the combination lowers the 10% effective-area energy threshold from 64 GeV to 30 GeV, improves the angular resolution by 57% at 100 GeV, and improves the detection sensitivity by 41% at the threshold, with the new analysis reproducing the Crab Nebula spectrum down to lower energies. The stakes are that the same monoscopic events are the ones that probe the faintest and most distant sources, and the methods transfer to mixed-size arrays such as the planned Cherenkov Telescope Array, where only the largest telescopes see low-energy events.","feed_headline":"Analysis upgrade halves the H.E.S.S. energy threshold","feed_subtitle":"A smarter image-orientation step and size-dependent selection cut also cut angular error by 57 percent.","key_machinery":"The load-bearing object is the 'flip': the binary decision of whether the true source position lies on the left or the right of the image's center of gravity in the camera frame, which fixes which end of the shower ellipse is the head. In the standard analysis the flip is read from the sign of the Hillas-skewness; here it is assigned by an XGBoost boosted decision tree whose most important input is the new kink-minor variable, the gradient of the image's off-axis position along the major axis, which encodes how the ellipse width changes from head to tail. The flip must be computed first because roughly half of the new parameters are asymmetry variables that swap or change sign under mirroring. The second piece of machinery is the size-dependent selection cut: the q-factor $\\epsilon_{\\rm sig}/\\sqrt{\\epsilon_{\\rm bkg}}$ is scanned in bins of total image intensity and the optimal cut is smoothed into a continuous curve, so low-size events are kept with a loose cut and large events with a strict one. The third piece is the parameter set itself: sector-based moments in head/center/tail sections, continuous gradient variables such as the profile-gradient, kink-major, and kink-minor, roughness and auto-correlation statistics, and brightest-pixel fractions and distances, all chosen by iterative variable-importance removal so that the final gamma-hadron separator uses 31 inputs.","core_discovery":"The central claim is that monoscopic gamma-ray events—showers seen by one telescope only, which are most of the events below about 100 GeV—are far better understood than the standard pipeline assumes. The paper argues that the standard reliance on the sign of the Hillas-skewness to decide the image orientation is the main limiting step: for small or truncated images the skewness is unreliable, and every downstream quantity (energy, direction, and the new head- and tail-dependent parameters) inherits the error. Feeding the image parameters to a trained classifier reduces the wrongly flipped fraction from 38% to 32% at threshold and from 29% to 17% at 100 GeV, and the angular resolution at 100 GeV improves by 57%. Making the gamma-hadron cut depend on image size, rather than a single constant cut, is the second independent gain: because separation quality rises with image size, a constant cut either discards useful small events or admits useless large ones, whereas a per-size cut keeps the q-factor near its optimum everywhere, which is what lowers the 10% effective-area threshold from 64 GeV to 30 GeV and the 10% energy-bias threshold from 120 GeV to 50 GeV. The new image parameters—sector moments, gradients along and across the ellipse, roughness, and brightest-pixel statistics—raise the q-factor by about 20% on average for low- and medium-size events, and together the three changes yield 41% better sensitivity at the low-energy threshold.","pith_inferences":["The two headline gains look separable: the size-dependent cut alone drops the energy threshold, while the new variables plus the flip-BDT carry most of the angular-resolution improvement. If so, either component could be ported to other single-telescope analyses without the other, which is useful for arrays whose cameras and trigger thresholds differ from H.E.S.S. CT5.","The paper's own decision to exclude time-based variables, because standard image cleaning leaves pixel times vulnerable to night-sky background, points to a concrete next step: a time-aware cleaning would likely let the time-gradient parameters into the separator and push the low-energy background rejection further.","The mirroring step that randomizes the flip distribution bakes in an assumption that performance is independent of source position in the camera; measuring whether the gains survive for sources at large offsets, where images are truncated most often, would test the method where it is supposedly most needed."],"forward_implications":["The 10% effective-area energy threshold falls from 64 GeV to 30 GeV and the 10% energy-bias threshold from 120 GeV to 50 GeV, so the telescope can claim credible detection and spectra roughly one octave lower in energy than before.","Angular resolution at 100 GeV improves by 57% and the false-flip fraction falls from 38% to 32% at threshold and from 29% to 17% at 100 GeV, which sharpens the separation of nearby sources and reduces background in aperture analysis.","Sensitivity for a 5-sigma detection in 50 hours improves by 41% at the low-energy threshold, and the q-factor of gamma-hadron separation rises on average by 20% for low- to medium-size events.","The new analysis reproduces the Crab Nebula spectrum measured by the previous chain and extends it to lower energies with a much smaller energy bias (41% versus 128% at 100 GeV), confirming the simulation-based expectations on real data.","The price of the gains is a modest rise in systematic sensitivity to night-sky background: the effective area changes by about 2.5% between nominal and 1.65-times-higher background, about one percentage point more than the standard analysis."],"supporting_citations":[{"why":"The standard H.E.S.S. monoscopic reconstruction, based on Hillas parameters and a constant selection cut; the baseline that this paper improves.","marker":"Murach et al. (2015)"},{"why":"Introduces the moment-based image parameters (the 'Hillas parameters') that form the foundation of the new variable set.","marker":"Hillas (1985)"},{"why":"Provides the calculation of the Hillas parameters that the paper references for the standard variables.","marker":"Reynolds et al. (1993)"},{"why":"CORSIKA, the air-shower simulation that generates the gamma-ray and proton events used for training and performance evaluation.","marker":"Heck et al. (1998)"},{"why":"sim_telarray, the telescope-response simulation used to model the CT5 camera with FlashCam.","marker":"Bernlöhr (2008)"},{"why":"XGBoost, the boosted-decision-tree package used for the flip determination and gamma-hadron separation.","marker":"Chen & Guestrin (2016)"},{"why":"Showed that separation performance improves with event size, motivating the size-dependent selection cut.","marker":"Voigt et al. (2014)"},{"why":"Demonstrated that pixel time information improves monoscopic point-source sensitivity for MAGIC, motivating the time-gradient parameters.","marker":"Aliu et al. (2009)"},{"why":"Provides the Crab Nebula mono dataset and reference spectrum used to validate the new analysis on real data.","marker":"Aharonian et al. (2024)"},{"why":"Reflected-region background estimation used in the Crab Nebula application.","marker":"Berge et al. (2007)"}],"fun_headline_variants":["H.E.S.S. upgrade halves low-energy threshold, boosts sensitivity 41%","Machine-learning orientation cuts H.E.S.S. angular error by 57%","Size-dependent cuts lower H.E.S.S. threshold from 64 to 30 GeV","New image parameters improve H.E.S.S. low-energy sensitivity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Every headline number is computed from simulated air showers and a simulated camera response that the paper assumes are faithful at low energies and across varying night-sky background—the paper keeps its preselection safely above trigger level because mis-modeling is a known risk there, and the Crab Nebula measurement is the only real-data check of this assumption.","fun_headline_variants_meta":{"raw":{"variants":["H.E.S.S. upgrade halves low-energy threshold, boosts sensitivity 41%","Machine-learning orientation cuts H.E.S.S. angular error by 57%","Size-dependent cuts lower H.E.S.S. threshold from 64 to 30 GeV","New image parameters improve H.E.S.S. low-energy sensitivity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000841,"raw_usage":{"total_tokens":3765,"prompt_tokens":1147,"completion_tokens":2618,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":763,"completion_tokens_details":{"reasoning_tokens":2535}},"tokens_in":763,"tokens_out":2618,"duration_ms":19897,"temperature":1.0,"reasoning_tokens":2535,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:19:45.236062+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the monoscopic Crab Nebula spectrum between 50 and 100 GeV with the new chain and compare it against the stereoscopic measurement of the same source: the paper claims an energy bias below 10% down to 50 GeV, so a mono spectrum departing from the stereo spectrum by more than the quoted systematic band in that range would show the simulated low-energy instrument response is wrong.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The standard H.E.S.S. monoscopic reconstruction, based on Hillas parameters and a constant selection cut; the baseline that this paper improves."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the moment-based image parameters (the 'Hillas parameters') that form the foundation of the new variable set."},{"cited_title":"T., Akerlof, C","cited_arxiv_id":null,"evidence_quote":"Provides the calculation of the Hillas parameters that the paper references for the standard variables."},{"cited_title":"1998, CORSIKA: A Monte Carlo code to simulate extensive air showers, Tech","cited_arxiv_id":null,"evidence_quote":"CORSIKA, the air-shower simulation that generates the gamma-ray and proton events used for training and performance evaluation."},{"cited_title":"A., et al","cited_arxiv_id":null,"evidence_quote":"Demonstrated that pixel time information improves monoscopic point-source sensitivity for MAGIC, motivating the time-gradient parameters."},{"cited_title":"A., Aschersleben, J., et al","cited_arxiv_id":null,"evidence_quote":"Provides the Crab Nebula mono dataset and reference spectrum used to validate the new analysis on real data."},{"cited_title":"2007, A&A, 466, 1219 Bernlöhr, K","cited_arxiv_id":null,"evidence_quote":"Reflected-region background estimation used in the Crab Nebula application."}],"review_version":1}