REVIEW 3 major objections 6 minor 25 references
HAWC Performance Enhanced by Machine Learning in Gamma-Hadron Separation
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A machine-learning classifier trained on 20 shower features separates gamma rays from hadrons in HAWC data better than the standard cut, lifting the Crab Nebula's significance from $299\sigma$ to $356\sigma$.
desk verdict A useful, real-data-validated ML upgrade for HAWC whose headline sensitivity gains are probably inflated by in-sample threshold optimization; the Crab gain is the more trustworthy claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a feedforward Multilayer Perceptron with two hidden layers (128 and 64 neurons), dropout, and a sigmoid output that scores each event's gamma-likelihood. It is trained once on Monte Carlo gamma-ray and cosmic-ray showers from 5 GeV to 500 TeV using binary cross-entropy loss and an adaptive-moment optimizer, then applied to real HAWC data. The 20 input features include the ten normalized annulus charge fractions, which encode the radial energy profile of the shower, and the classifier's score is turned into event selection by scanning the gamma-probability threshold to maximize the Q-factor, $\xi_\gamma/\sqrt{\xi_h}$, in each fHit energy bin (fHit being the fraction of hit channels, a proxy for shower energy).
What would settle it
Recompute the Crab analysis using a background sample built only from sky regions with no known gamma-ray sources, or with all bright sources masked; if the MLP's 19% significance gain over the Standard Cut disappears, part of the improvement is an artifact of gamma-ray contamination in the real-data background.
Extended reading notes
Core claim
An MLP classifier operating on 20 features from HAWC's Pass 5 reconstruction—including the fraction of hit channels, two energy estimators, lateral-distribution fit quality, PINCness, compactness, and ten concentric annulus charge fractions—outperforms the Standard Cut in gamma-hadron separation. On real data, the model raises the Crab Nebula's peak significance from $299\sigma$ to $356\sigma$, a 19% enhancement, and improves the differential sensitivity by roughly 23% for on-array events (shower cores inside the main array) and 40% for off-array events (cores within 150 m of the array center). The gain is concentrated in the 1–20 TeV range; at high energies the muon content of hadronic showers already makes separation easy, and at low energies the showers are intrinsically similar. The MLP also lifts the Galactic Center significance by 34.3% and Boomerang by 5.3% while keeping the Crab's spectral energy distribution consistent with the standard analysis.
Load-bearing premise
The whole result rests on simulated gamma-ray and cosmic-ray showers matching what HAWC actually records after reweighting, and on the real-data background sample containing almost no gamma rays from known sources.
Editorial extensions
If this is right
- HAWC's standard pipeline can adopt the MLP score in place of the Standard Cut and gain 19% significance on the Crab, with per-bin gains up to 53% in the mid-energy off-array bin.
- The sensitivity improvement is strongest at 1–20 TeV, where combining many variables matters most, making weak sources in this range easier to detect.
- Off-array events show a 40% sensitivity gain, effectively enlarging the usable detector area for gamma-ray astronomy.
- The Crab's measured spectral energy distribution is unchanged under the MLP selection, so the added sensitivity does not come with spectral bias.
- The same 20-feature approach is presented as a reference for other ground-based observatories planning water-Cherenkov or air-shower detection.
Reading between the lines
- Because the background proxy is real HAWC data rather than a gamma-ray-free sample, the reported Q-factors and significance gains may be inflated by gamma rays from bright sources in the field; a source-masked re-analysis would bound this effect.
- The model's lack of explicit declination conditioning at high energies likely explains why the MLP advantage shrinks above tens of TeV; a version conditioned on zenith or declination might extend the gain.
- The training relies on Monte Carlo feature distributions matching real data after reweighting; a direct comparison of fHit and annulus-charge distributions in a gamma-ray-free control region would test whether the simulated training sample is trustworthy.
- The unified single-model design suggests the approach transfers to other arrays, but the optimal feature set and threshold scan would need retuning for each detector geometry.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a machine-learning-based gamma-hadron separation for the HAWC Observatory using the Pass 5 reconstruction. The authors train three classifiers (BDT, MLP, CNN) on Monte Carlo gamma-ray and hadron showers using 20 input features and compare them with the standard cut (SC) using a Q-factor metric in which gamma efficiency is estimated from simulation and hadron efficiency from a three-year real-data background sample. The MLP is reported as the best method, yielding about 23% (on-array) and 40% (off-array) improvements in differential sensitivity and a 19% increase in the peak significance of the Crab Nebula (356 sigma vs 299 sigma), together with consistent spectral energy distributions, plus significance gains for the Galactic Center and Boomerang.
Significance. If the reported improvements are robust, the work is of practical importance: it demonstrates on real data that a modern ML classifier can modestly improve the sensitivity of an operating water-Cherenkov observatory, and it offers a template for LHAASO and SWGO. The paper's strengths include a 10-year real-data validation on multiple sources, spectral consistency checks, a comparison of several ML architectures, and explicit presentation of the Q-factor and sensitivity definitions. Its main limitations are the absence of an independent validation split in the threshold optimization and the lack of a quantitative MC-to-data closure test, so the numerical claims should be treated as preliminary until these are addressed.
major comments (3)
- [4.6 and 5.1] The p_gamma threshold per fHit bin is selected by maximizing the Q-factor on the same three-year real-data background sample that is later used to compute the sensitivity curves (Section 3.2 explicitly says this sample is used for 'evaluating the gamma-hadron separation performance and sensitivity curves'). No split into optimization and evaluation samples, no cross-validation, and no trial-factor correction is described. Because the scan is over a continuous threshold and the background sample retains only 'hundreds' of high-energy events, this in-sample optimization can inflate the Q-factor and therefore the reported sensitivity improvements (23% on-array, 40% off-array) and the Crab peak-significance gain, given that the 10-year source dataset overlaps the tuning period. The authors should optimize on one subset, evaluate on a disjoint subset, and report the threshold variance or a trial-factor-corrected Q.
- [3.1-3.2 and Figures 1-3] The Q-factor evaluation combines Monte-Carlo gamma efficiencies with real-data hadron efficiencies, but the paper does not show a quantitative verification that the reconstructed feature distributions (fHit, fAnnulusQ0-fAnnulusQ9, NNEnergy/GPEnergy, LDF variables) agree between simulation and data after the standard reweighting. If the simulation mismodels the gamma-ray event properties, the gamma efficiency and the position of the optimal p_gamma threshold will be biased, and the source-level checks in Section 5 are only indirect evidence of transferability. A closure test using a gamma-ray-enriched real-data sample (e.g., the Crab excess or the Moon shadow) should be added, and its associated systematic uncertainty should be propagated to the sensitivity curves.
- [5.2-5.3] The headline significance improvements for Crab, Galactic Center, and Boomerang are quoted without statistical uncertainties on the relative gains and without any statement about whether these sources were selected before applying the MLP analysis. If these sources were chosen after examining MLP results over a broader source list, the 34.3% and 5.3% gains are subject to a look-elsewhere effect. At minimum, the paper should state that the sources were preselected for the stated declination-dependence test and give a bootstrap or profile-likelihood uncertainty on the (sigma_MLP - sigma_SC)/sigma_SC ratios.
minor comments (6)
- [4.4] The sentence 'This a value chosen after testing several options' should read 'This value was chosen after testing several options'.
- [Table 1] The B0C0 bin shows a -0.75% change; the text acknowledges this, but the abstract and conclusion should explicitly qualify the 19% claim as referring to the peak significance over the full energy range, not to all bins.
- [5.3] The quoted SC significances for Galactic Center (6.43 sigma) and Boomerang (14.40 sigma) come from previous papers; the authors should confirm that the same livetime, event selection, and energy range are used in the MLP comparison, or the 34.3% and 5.3% gains are not apples-to-apples.
- [4.6, Eq. (1)] The assumption of a 'sufficiently large background' should be quantified (e.g., the number of background events per fHit bin after the cut), since the Q-factor formula's validity depends on it.
- [Figure 3 caption] The phrase 'A solid ROC curve lying below the gray dashed line' is ambiguous; it should read 'A ROC curve lying below the gray dashed line'.
- [3.2] The 3-year background sample selection is described only by fHit filtering; the authors should specify whether known gamma-ray sources were masked, since bright sources in the field of view would contaminate the hadron-efficiency estimate (in a conservative direction, but still affecting the Q-factor).
Circularity Check
Sensitivity gain is partly an in-sample fit: the pγ threshold is optimized on the same 3-year real-data background sample used to compute the reported sensitivity curves, making the 23%/40% improvements statistically forced rather than out-of-sample predictions.
-
fitted input called prediction
[Sections 3.2 and 4.6; used in Section 5.1]
"We selected a 3-year dataset as the background sample for evaluating the gamma-hadron separation performance and sensitivity curves, as described in Section 4.6 and Section 5.1. ... The optimal Q-factor search involves scanning the pγ value from 0 to 1 ... search the best Q-factor by scanning the pγ cut."
The pγ threshold is a fitted parameter, selected by maximizing the Q-factor on the 3-year real-data background sample. The same sample is then used to compute the sensitivity curves that yield the headline improvements of ~23% (on-array) and ~40% (off-array). Since sensitivity is directly tied to Q (higher Q gives better flux sensitivity at fixed gamma efficiency), optimizing Q on the evaluation sample picks the threshold that makes that sample look best. The reported gain over SC is therefore not an independent measurement; it is the result of in-sample optimization. The Crab significance improvement (19%) uses a larger 3070-day dataset, but the thresholds were selected on the overlapping 3-year subset, so the tuning bias partially persists.
full rationale
The paper's core activity is an empirical performance comparison, not a mathematical derivation. The MLP itself is trained on Monte Carlo data with true labels, so the classifier's ranking ability is not defined in terms of the real-data result and has independent support from MC ROC/AUC curves. However, the evaluation protocol is partially circular in the sense of pattern 2: the operating point (pγ cut) is tuned to maximize Q-factor on the very real-data background sample that is then used to compute sensitivity and, indirectly, the source-significance improvements. With only 'at least hundreds' of high-energy events in the background sample, scanning a continuous threshold on the same events can exploit statistical fluctuations, inflating the Q-factor and the derived sensitivity gain. This does not invalidate the possibility that MLP genuinely improves separation, but it means the reported 23%/40%/19% numbers are not clean out-of-sample predictions. Other self-citations (Pass 5, previous HAWC analyses) are not load-bearing; they merely describe the baseline and prior feature definitions. No uniqueness theorems or ansatz-smuggling circularity are present. The severity is moderate: the central architecture comparison has independent content, but the headline quantitative gains are partly an artifact of in-sample threshold selection.
Assumptions & free parameters
free parameters (4)
- MLP hidden layer sizes, dropout, learning rate =
128 and 64 neurons, 20% dropout, learning rate 0.01
- BDT hyperparameters =
500 trees, depth 6, learning rate 0.2, 50% bagging, 20 cuts
- p_gamma threshold per fHit bin =
Scanned 0 to 1; value chosen to maximize Q-factor
- fHit bin boundaries and low-fHit cut =
B0-B10 bins; low-fHit events excluded for the background sample
assumptions (4)
- domain assumption CORSIKA with QGSJET-II-04 and FLUKA, combined with GEANT4 simulation of HAWC, accurately reproduces real air-shower and detector response.
- domain assumption Real HAWC data can serve as a hadron background sample for evaluation without significant gamma contamination.
- domain assumption The Q-factor metric with the large-background approximation is a valid proxy for significance improvement.
- domain assumption Simulated spectra and core weights can be reweighted to preserve true physical distributions.
Cite this review
Pith. "Pith review of HAWC Performance Enhanced by Machine Learning in Gamma-Hadron Separation." pith.science (2026). https://pith.science/paper/ZKL7YV6L
@misc{pith2026250618277,
author = {Pith},
title = {Pith review of: HAWC Performance Enhanced by Machine Learning in Gamma-Hadron Separation},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZKL7YV6L}},
note = {Machine review of arXiv:2506.18277}
}
read the original abstract
Improving gamma-hadron separation is one of the most effective ways to enhance the performance of ground-based gamma-ray observatories. With over a decade of continuous operation, the High-Altitude Water Cherenkov (HAWC) Observatory has contributed significantly to high-energy astrophysics. To further leverage its rich dataset, we introduce a machine learning approach for gamma-hadron separation. A Multilayer Perceptron shows the best performance, surpassing traditional and other Machine Learning based methods. This approach shows a notable improvement in the detector's sensitivity, supported by results from both simulated and real HAWC data. In particular, it achieves a 19\% increase in significance for the Crab Nebula, commonly used as a benchmark. These improvements highlight the potential of machine learning to significantly enhance the performance of HAWC and provide a valuable reference for ground-based observatories, such as Large High Altitude Air Shower Observatory (LHAASO) and the upcoming Southern Wide-field Gamma-ray Observatory (SWGO).
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
2017, The Astrophysical Journal, 843, 39
Abeysekara, A., Albert, A., Alfaro, R., et al. 2017, The Astrophysical Journal, 843, 39
work page 2017
-
[2]
2019, The Astrophysical Journal, 881, 134
Abeysekara, A., Albert, A., Alfaro, R., et al. 2019, The Astrophysical Journal, 881, 134
work page 2019
-
[3]
2023, Nuclear Instruments and Methods in Physics Research Section A:
Abeysekara, A., Albert, A., Alfaro, R., et al. 2023, Nuclear Instruments and Methods in Physics Research Section A:
2023
-
[4]
2021, Chinese Physics C, 45, 025002
Aharonian, F., An, Q., Bai, L., et al. 2021, Chinese Physics C, 45, 025002
work page 2021
-
[5]
2008, Nuclear Instruments and Methods in Physics Research Section A:
Albert, J., Aliu, E., Anderhub, H., et al. 2008, Nuclear Instruments and Methods in Physics Research Section A:
work page 2008
-
[6]
2022, Nuclear Instruments and Methods in Physics Research Section A:
Alfaro, R., Alvarez, C., ´Alvarez, J., et al. 2022, Nuclear Instruments and Methods in Physics Research Section A:
work page 2022
-
[7]
2024, Astronomy & Astrophysics, 691, A89
Alfaro, R., Alvarez, C., Arteaga-Vel´ azquez, J., et al. 2024, Astronomy & Astrophysics, 691, A89
work page 2024
- [8]
Show all 25 references
-
[9]
2003, The Astrophysical Journal, 595, 803
Atkins, R., Benbow, W., Berley, D., et al. 2003, The Astrophysical Journal, 595, 803
2003
-
[10]
2019, arXiv e-prints, arXiv
Bai, X., Bi, B., Bi, X., et al. 2019, arXiv e-prints, arXiv
2019
-
[11]
1997–2024,, urlhttps://root.cern/ Collaboration*†, L., Cao, Z., Aharonian, F., et al
Brun, R., & Rademakers, F. 1997–2024,, urlhttps://root.cern/ Collaboration*†, L., Cao, Z., Aharonian, F., et al. 2021, Science, 373, 425
1997
-
[12]
2006, Pattern recognition letters, 27, 861
Fawcett, T. 2006, Pattern recognition letters, 27, 861
2006
-
[13]
1998, Report fzka, 6019 H¨ ocker, A., Speckmayer, P., Tegenfeldt, F., Stelzer, J., &
Heck, D., Knapp, J., Capdevielle, J., et al. 1998, Report fzka, 6019 H¨ ocker, A., Speckmayer, P., Tegenfeldt, F., Stelzer, J., &
1998
-
[14]
1958, Progress of Theoretical Physics Supplement, 6, 93
Kamata, K., & Nishimura, J. 1958, Progress of Theoretical Physics Supplement, 6, 93
1958
- [15]
-
[16]
2025, arXiv preprint arXiv:2501.19064
Kryukov, A., Demichev, A., & Ilyin, V. 2025, arXiv preprint arXiv:2501.19064
2025 arXiv
-
[17]
2025, The Astrophysical Journal Supplement Series, 276, 24
Li, J., Lv, H., Liu, Y., et al. 2025, The Astrophysical Journal Supplement Series, 276, 24
2025
-
[18]
2019, arXiv preprint arXiv:1908.07634
Marandon, V., Jardin-Blicq, A., & Schoorlemmer, H. 2019, arXiv preprint arXiv:1908.07634
2019 arXiv
-
[19]
2009, Astroparticle Physics, 31, 383
Ohm, S., van Eldik, C., & Egberts, K. 2009, Astroparticle Physics, 31, 383
2009
-
[20]
2023, Nuclear Instruments and Methods in Physics Research Section A: Accelerators, Spectrometers, Detectors and Associated Equipment, 1055, 168442
Ohm, S., Wagner, S., Collaboration, H., et al. 2023, Nuclear Instruments and Methods in Physics Research Section A: Accelerators, Spectrometers, Detectors and Associated Equipment, 1055, 168442
2023
-
[21]
D., & D’Anna, F
Pagliaro, A., Staiti, G. D., & D’Anna, F. 2011, Nuclear Physics B-Proceedings Supplements, 212, 286
2011
-
[22]
2019, in Advances in neural information processing systems, 8024–8035
Paszke, A., Gross, S., Massa, F., et al. 2019, in Advances in neural information processing systems, 8024–8035
2019
-
[23]
2018, Astroparticle Physics, 99, 43
Tian, Z., Wang, Z., Liu, Y., et al. 2018, Astroparticle Physics, 99, 43
2018
-
[24]
2019, in 36th International Cosmic Ray Conference (ICRC2019), Vol
Wang, X. 2019, in 36th International Cosmic Ray Conference (ICRC2019), Vol. 36, 820
2019
-
[25]
1995, Astroparticle Physics, 4, 119
Westerhoff, S., Funk, B., Lindner, A., et al. 1995, Astroparticle Physics, 4, 119
1995
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.