Pith. sign in

REVIEW 3 major objections 6 minor 25 references

HAWC Performance Enhanced by Machine Learning in Gamma-Hadron Separation

T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A machine-learning classifier trained on 20 shower features separates gamma rays from hadrons in HAWC data better than the standard cut, lifting the Crab Nebula's significance from $299\sigma$ to $356\sigma$.

desk verdict A useful, real-data-validated ML upgrade for HAWC whose headline sensitivity gains are probably inflated by in-sample threshold optimization; the Crab gain is the more trustworthy claim. read the letter →

arxiv 2506.18277 v1 pith:ZKL7YV6L submitted 2025-06-23 astro-ph.IM astro-ph.HE

classification astro-ph.IMastro-ph.HE
keywords gamma-hadronseparationmultilayerperceptronHAWCobservatorywaterCherenkovdetectorvery-high-energygammaraysairshowersQ-factormachinelearningclassification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Ground-based gamma-ray observatories are swamped by cosmic-ray air showers that mimic gamma-ray signals. This paper claims that a Multilayer Perceptron trained on 20 reconstructed event features across the full energy range separates gamma-ray from hadron showers more cleanly than the current standard cut. The payoff is concrete: for the Crab Nebula, a standard benchmark, the detection significance rises from $299\sigma$ to $356\sigma$, a 19% increase, and differential sensitivity improves by about 23% for on-array and 40% for off-array events. The authors argue this makes weak and mid-energy TeV sources easier to detect and provides a template for other ground-based water-Cherenkov and air-shower arrays.

What carries the argument

The central object is a feedforward Multilayer Perceptron with two hidden layers (128 and 64 neurons), dropout, and a sigmoid output that scores each event's gamma-likelihood. It is trained once on Monte Carlo gamma-ray and cosmic-ray showers from 5 GeV to 500 TeV using binary cross-entropy loss and an adaptive-moment optimizer, then applied to real HAWC data. The 20 input features include the ten normalized annulus charge fractions, which encode the radial energy profile of the shower, and the classifier's score is turned into event selection by scanning the gamma-probability threshold to maximize the Q-factor, $\xi_\gamma/\sqrt{\xi_h}$, in each fHit energy bin (fHit being the fraction of hit channels, a proxy for shower energy).

What would settle it

Recompute the Crab analysis using a background sample built only from sky regions with no known gamma-ray sources, or with all bright sources masked; if the MLP's 19% significance gain over the Standard Cut disappears, part of the improvement is an artifact of gamma-ray contamination in the real-data background.

Watch

Extended reading notes

Core claim

An MLP classifier operating on 20 features from HAWC's Pass 5 reconstruction—including the fraction of hit channels, two energy estimators, lateral-distribution fit quality, PINCness, compactness, and ten concentric annulus charge fractions—outperforms the Standard Cut in gamma-hadron separation. On real data, the model raises the Crab Nebula's peak significance from $299\sigma$ to $356\sigma$, a 19% enhancement, and improves the differential sensitivity by roughly 23% for on-array events (shower cores inside the main array) and 40% for off-array events (cores within 150 m of the array center). The gain is concentrated in the 1–20 TeV range; at high energies the muon content of hadronic showers already makes separation easy, and at low energies the showers are intrinsically similar. The MLP also lifts the Galactic Center significance by 34.3% and Boomerang by 5.3% while keeping the Crab's spectral energy distribution consistent with the standard analysis.

Load-bearing premise

The whole result rests on simulated gamma-ray and cosmic-ray showers matching what HAWC actually records after reweighting, and on the real-data background sample containing almost no gamma rays from known sources.

Editorial extensions

If this is right

  • HAWC's standard pipeline can adopt the MLP score in place of the Standard Cut and gain 19% significance on the Crab, with per-bin gains up to 53% in the mid-energy off-array bin.
  • The sensitivity improvement is strongest at 1–20 TeV, where combining many variables matters most, making weak sources in this range easier to detect.
  • Off-array events show a 40% sensitivity gain, effectively enlarging the usable detector area for gamma-ray astronomy.
  • The Crab's measured spectral energy distribution is unchanged under the MLP selection, so the added sensitivity does not come with spectral bias.
  • The same 20-feature approach is presented as a reference for other ground-based observatories planning water-Cherenkov or air-shower detection.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the background proxy is real HAWC data rather than a gamma-ray-free sample, the reported Q-factors and significance gains may be inflated by gamma rays from bright sources in the field; a source-masked re-analysis would bound this effect.
  • The model's lack of explicit declination conditioning at high energies likely explains why the MLP advantage shrinks above tens of TeV; a version conditioned on zenith or declination might extend the gain.
  • The training relies on Monte Carlo feature distributions matching real data after reweighting; a direct comparison of fHit and annulus-charge distributions in a gamma-ray-free control region would test whether the simulated training sample is trustworthy.
  • The unified single-model design suggests the approach transfers to other arrays, but the optimal feature set and threshold scan would need retuning for each detector geometry.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper presents a machine-learning-based gamma-hadron separation for the HAWC Observatory using the Pass 5 reconstruction. The authors train three classifiers (BDT, MLP, CNN) on Monte Carlo gamma-ray and hadron showers using 20 input features and compare them with the standard cut (SC) using a Q-factor metric in which gamma efficiency is estimated from simulation and hadron efficiency from a three-year real-data background sample. The MLP is reported as the best method, yielding about 23% (on-array) and 40% (off-array) improvements in differential sensitivity and a 19% increase in the peak significance of the Crab Nebula (356 sigma vs 299 sigma), together with consistent spectral energy distributions, plus significance gains for the Galactic Center and Boomerang.

Significance. If the reported improvements are robust, the work is of practical importance: it demonstrates on real data that a modern ML classifier can modestly improve the sensitivity of an operating water-Cherenkov observatory, and it offers a template for LHAASO and SWGO. The paper's strengths include a 10-year real-data validation on multiple sources, spectral consistency checks, a comparison of several ML architectures, and explicit presentation of the Q-factor and sensitivity definitions. Its main limitations are the absence of an independent validation split in the threshold optimization and the lack of a quantitative MC-to-data closure test, so the numerical claims should be treated as preliminary until these are addressed.

major comments (3)
  1. [4.6 and 5.1] The p_gamma threshold per fHit bin is selected by maximizing the Q-factor on the same three-year real-data background sample that is later used to compute the sensitivity curves (Section 3.2 explicitly says this sample is used for 'evaluating the gamma-hadron separation performance and sensitivity curves'). No split into optimization and evaluation samples, no cross-validation, and no trial-factor correction is described. Because the scan is over a continuous threshold and the background sample retains only 'hundreds' of high-energy events, this in-sample optimization can inflate the Q-factor and therefore the reported sensitivity improvements (23% on-array, 40% off-array) and the Crab peak-significance gain, given that the 10-year source dataset overlaps the tuning period. The authors should optimize on one subset, evaluate on a disjoint subset, and report the threshold variance or a trial-factor-corrected Q.
  2. [3.1-3.2 and Figures 1-3] The Q-factor evaluation combines Monte-Carlo gamma efficiencies with real-data hadron efficiencies, but the paper does not show a quantitative verification that the reconstructed feature distributions (fHit, fAnnulusQ0-fAnnulusQ9, NNEnergy/GPEnergy, LDF variables) agree between simulation and data after the standard reweighting. If the simulation mismodels the gamma-ray event properties, the gamma efficiency and the position of the optimal p_gamma threshold will be biased, and the source-level checks in Section 5 are only indirect evidence of transferability. A closure test using a gamma-ray-enriched real-data sample (e.g., the Crab excess or the Moon shadow) should be added, and its associated systematic uncertainty should be propagated to the sensitivity curves.
  3. [5.2-5.3] The headline significance improvements for Crab, Galactic Center, and Boomerang are quoted without statistical uncertainties on the relative gains and without any statement about whether these sources were selected before applying the MLP analysis. If these sources were chosen after examining MLP results over a broader source list, the 34.3% and 5.3% gains are subject to a look-elsewhere effect. At minimum, the paper should state that the sources were preselected for the stated declination-dependence test and give a bootstrap or profile-likelihood uncertainty on the (sigma_MLP - sigma_SC)/sigma_SC ratios.
minor comments (6)
  1. [4.4] The sentence 'This a value chosen after testing several options' should read 'This value was chosen after testing several options'.
  2. [Table 1] The B0C0 bin shows a -0.75% change; the text acknowledges this, but the abstract and conclusion should explicitly qualify the 19% claim as referring to the peak significance over the full energy range, not to all bins.
  3. [5.3] The quoted SC significances for Galactic Center (6.43 sigma) and Boomerang (14.40 sigma) come from previous papers; the authors should confirm that the same livetime, event selection, and energy range are used in the MLP comparison, or the 34.3% and 5.3% gains are not apples-to-apples.
  4. [4.6, Eq. (1)] The assumption of a 'sufficiently large background' should be quantified (e.g., the number of background events per fHit bin after the cut), since the Q-factor formula's validity depends on it.
  5. [Figure 3 caption] The phrase 'A solid ROC curve lying below the gray dashed line' is ambiguous; it should read 'A ROC curve lying below the gray dashed line'.
  6. [3.2] The 3-year background sample selection is described only by fHit filtering; the authors should specify whether known gamma-ray sources were masked, since bright sources in the field of view would contaminate the hadron-efficiency estimate (in a conservative direction, but still affecting the Q-factor).

Circularity Check

1 steps flagged · score 5.0 of 10

Sensitivity gain is partly an in-sample fit: the pγ threshold is optimized on the same 3-year real-data background sample used to compute the reported sensitivity curves, making the 23%/40% improvements statistically forced rather than out-of-sample predictions.

  1. fitted input called prediction [Sections 3.2 and 4.6; used in Section 5.1]
    "We selected a 3-year dataset as the background sample for evaluating the gamma-hadron separation performance and sensitivity curves, as described in Section 4.6 and Section 5.1. ... The optimal Q-factor search involves scanning the pγ value from 0 to 1 ... search the best Q-factor by scanning the pγ cut."

    The pγ threshold is a fitted parameter, selected by maximizing the Q-factor on the 3-year real-data background sample. The same sample is then used to compute the sensitivity curves that yield the headline improvements of ~23% (on-array) and ~40% (off-array). Since sensitivity is directly tied to Q (higher Q gives better flux sensitivity at fixed gamma efficiency), optimizing Q on the evaluation sample picks the threshold that makes that sample look best. The reported gain over SC is therefore not an independent measurement; it is the result of in-sample optimization. The Crab significance improvement (19%) uses a larger 3070-day dataset, but the thresholds were selected on the overlapping 3-year subset, so the tuning bias partially persists.

full rationale

The paper's core activity is an empirical performance comparison, not a mathematical derivation. The MLP itself is trained on Monte Carlo data with true labels, so the classifier's ranking ability is not defined in terms of the real-data result and has independent support from MC ROC/AUC curves. However, the evaluation protocol is partially circular in the sense of pattern 2: the operating point (pγ cut) is tuned to maximize Q-factor on the very real-data background sample that is then used to compute sensitivity and, indirectly, the source-significance improvements. With only 'at least hundreds' of high-energy events in the background sample, scanning a continuous threshold on the same events can exploit statistical fluctuations, inflating the Q-factor and the derived sensitivity gain. This does not invalidate the possibility that MLP genuinely improves separation, but it means the reported 23%/40%/19% numbers are not clean out-of-sample predictions. Other self-citations (Pass 5, previous HAWC analyses) are not load-bearing; they merely describe the baseline and prior feature definitions. No uniqueness theorems or ansatz-smuggling circularity are present. The severity is moderate: the central architecture comparison has independent content, but the headline quantitative gains are partly an artifact of in-sample threshold selection.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central claim rests on simulated air-shower realism, the use of real data as a background proxy, and the choice of operating thresholds. No new physical entities are postulated.

free parameters (4)
  • MLP hidden layer sizes, dropout, learning rate = 128 and 64 neurons, 20% dropout, learning rate 0.01
    Chosen empirically; classification performance depends on these values (Section 4.4).
  • BDT hyperparameters = 500 trees, depth 6, learning rate 0.2, 50% bagging, 20 cuts
    Selected by parameter scans and validation (Section 4.3).
  • p_gamma threshold per fHit bin = Scanned 0 to 1; value chosen to maximize Q-factor
    Operating point selection affects the quoted efficiencies and sensitivities (Section 4.6).
  • fHit bin boundaries and low-fHit cut = B0-B10 bins; low-fHit events excluded for the background sample
    Analysis convention; sensitivity and significance are reported per bin (Sections 3.2 and 5.2).
assumptions (4)
  • domain assumption CORSIKA with QGSJET-II-04 and FLUKA, combined with GEANT4 simulation of HAWC, accurately reproduces real air-shower and detector response.
    Training labels and Q-factor evaluation depend on simulated gamma and hadron events matching real data (Section 3.1).
  • domain assumption Real HAWC data can serve as a hadron background sample for evaluation without significant gamma contamination.
    Section 3.2 uses real data as background; source photons from Crab or other objects could bias background estimates if not subtracted.
  • domain assumption The Q-factor metric with the large-background approximation is a valid proxy for significance improvement.
    Equation 1 assumes sufficiently large background; the paper asserts this is validated (Section 4.6).
  • domain assumption Simulated spectra and core weights can be reweighted to preserve true physical distributions.
    Training relies on reweighted MC events; incorrect reweighting would bias the classifier (Section 3.1).

how reviews work

0 comments
Cite this review

Pith. "Pith review of HAWC Performance Enhanced by Machine Learning in Gamma-Hadron Separation." pith.science (2026). https://pith.science/paper/ZKL7YV6L

@misc{pith2026250618277,
  author       = {Pith},
  title        = {Pith review of: HAWC Performance Enhanced by Machine Learning in Gamma-Hadron Separation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZKL7YV6L}},
  note         = {Machine review of arXiv:2506.18277}
}
read the original abstract

Improving gamma-hadron separation is one of the most effective ways to enhance the performance of ground-based gamma-ray observatories. With over a decade of continuous operation, the High-Altitude Water Cherenkov (HAWC) Observatory has contributed significantly to high-energy astrophysics. To further leverage its rich dataset, we introduce a machine learning approach for gamma-hadron separation. A Multilayer Perceptron shows the best performance, surpassing traditional and other Machine Learning based methods. This approach shows a notable improvement in the detector's sensitivity, supported by results from both simulated and real HAWC data. In particular, it achieves a 19\% increase in significance for the Crab Nebula, commonly used as a benchmark. These improvements highlight the potential of machine learning to significantly enhance the performance of HAWC and provide a valuable reference for ground-based observatories, such as Large High Altitude Air Shower Observatory (LHAASO) and the upcoming Southern Wide-field Gamma-ray Observatory (SWGO).

Figures

Figures reproduced from arXiv: 2506.18277 by the authors.

Figure 1
Figure 1. Gamma-ray efficiencies (dashed lines) and hadron efficiencies (solid lines) as a function of fHit bins for different classification methods. (a): Results for on-array events. (b): Results for off-array events. Classification methods compared include SC, MLP, CNN, and BDT [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Q-factor as a function of fHit bins for different classification methods. (a): Results for on-array events. (b): Results for off-array events. Classification methods compared include SC, MLP, CNN, and BDT. Q-factors improve with increasing fHit bins, with notable differences across methods, especially in high-fHit bins. classification thresholds, helping to assess how well the model distinguishes gamma-ray from hadr… view at source ↗
Figure 3
Figure 3. Receiver Operating Characteristic (ROC) curves and Area Under the Curve (AUC) values for MLP model perfor￾mance across different fHit bins. Left panel (a): Results for on-array events. Right panel (b): Results for off-array events. Each curve corresponds to a specific fHit bin, labeled from B0 to B10. The MLP model exhibits progressively better classification performance with increasing fHit, as indicated by higher … view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: HAWC 10-year sensitivity comparison between MLP and SC methods. Top: Comparison of differential sensitivity curves between the MLP-based event selection (solid lines) and the SC method (dashed lines) as a function of reconstructed gamma-ray energy. Bottom: Relative imp…
Figure 5
Figure 5. Figure 5: Spectral energy distribution of the Crab Nebula. The plot compares different analysis methods and datasets, including Pass 5 fHit bins with ML (black line), Pass 5 fHit bins with SC (red line), Pass 5 NN with SC (green line) and Pass 5 NN ML (blue line). The HAWC Pass …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

25 extracted references · 22 canonical work pages

  1. [1]

    2017, The Astrophysical Journal, 843, 39

    Abeysekara, A., Albert, A., Alfaro, R., et al. 2017, The Astrophysical Journal, 843, 39

  2. [2]

    2019, The Astrophysical Journal, 881, 134

    Abeysekara, A., Albert, A., Alfaro, R., et al. 2019, The Astrophysical Journal, 881, 134

  3. [3]

    2023, Nuclear Instruments and Methods in Physics Research Section A:

    Abeysekara, A., Albert, A., Alfaro, R., et al. 2023, Nuclear Instruments and Methods in Physics Research Section A:

  4. [4]

    2021, Chinese Physics C, 45, 025002

    Aharonian, F., An, Q., Bai, L., et al. 2021, Chinese Physics C, 45, 025002

  5. [5]

    2008, Nuclear Instruments and Methods in Physics Research Section A:

    Albert, J., Aliu, E., Anderhub, H., et al. 2008, Nuclear Instruments and Methods in Physics Research Section A:

  6. [6]

    2022, Nuclear Instruments and Methods in Physics Research Section A:

    Alfaro, R., Alvarez, C., ´Alvarez, J., et al. 2022, Nuclear Instruments and Methods in Physics Research Section A:

  7. [7]

    2024, Astronomy & Astrophysics, 691, A89

    Alfaro, R., Alvarez, C., Arteaga-Vel´ azquez, J., et al. 2024, Astronomy & Astrophysics, 691, A89

  8. [8]

    2007, Nuclear Physics News, 17, 20

    Allison, J. 2007, Nuclear Physics News, 17, 20

Show all 25 references
  1. [9]

    2003, The Astrophysical Journal, 595, 803

    Atkins, R., Benbow, W., Berley, D., et al. 2003, The Astrophysical Journal, 595, 803

  2. [10]

    2019, arXiv e-prints, arXiv

    Bai, X., Bi, B., Bi, X., et al. 2019, arXiv e-prints, arXiv

  3. [11]

    1997–2024,, urlhttps://root.cern/ Collaboration*†, L., Cao, Z., Aharonian, F., et al

    Brun, R., & Rademakers, F. 1997–2024,, urlhttps://root.cern/ Collaboration*†, L., Cao, Z., Aharonian, F., et al. 2021, Science, 373, 425

  4. [12]

    2006, Pattern recognition letters, 27, 861

    Fawcett, T. 2006, Pattern recognition letters, 27, 861

  5. [13]

    1998, Report fzka, 6019 H¨ ocker, A., Speckmayer, P., Tegenfeldt, F., Stelzer, J., &

    Heck, D., Knapp, J., Capdevielle, J., et al. 1998, Report fzka, 6019 H¨ ocker, A., Speckmayer, P., Tegenfeldt, F., Stelzer, J., &

  6. [14]

    1958, Progress of Theoretical Physics Supplement, 6, 93

    Kamata, K., & Nishimura, J. 1958, Progress of Theoretical Physics Supplement, 6, 93

  7. [15]

    P., & Ba, J

    Kingma, D. P., & Ba, J. 2014, arXiv preprint arXiv:1412.6980

  8. [16]

    2025, arXiv preprint arXiv:2501.19064

    Kryukov, A., Demichev, A., & Ilyin, V. 2025, arXiv preprint arXiv:2501.19064

  9. [17]

    2025, The Astrophysical Journal Supplement Series, 276, 24

    Li, J., Lv, H., Liu, Y., et al. 2025, The Astrophysical Journal Supplement Series, 276, 24

  10. [18]

    2019, arXiv preprint arXiv:1908.07634

    Marandon, V., Jardin-Blicq, A., & Schoorlemmer, H. 2019, arXiv preprint arXiv:1908.07634

  11. [19]

    2009, Astroparticle Physics, 31, 383

    Ohm, S., van Eldik, C., & Egberts, K. 2009, Astroparticle Physics, 31, 383

  12. [20]

    2023, Nuclear Instruments and Methods in Physics Research Section A: Accelerators, Spectrometers, Detectors and Associated Equipment, 1055, 168442

    Ohm, S., Wagner, S., Collaboration, H., et al. 2023, Nuclear Instruments and Methods in Physics Research Section A: Accelerators, Spectrometers, Detectors and Associated Equipment, 1055, 168442

  13. [21]

    D., & D’Anna, F

    Pagliaro, A., Staiti, G. D., & D’Anna, F. 2011, Nuclear Physics B-Proceedings Supplements, 212, 286

  14. [22]

    2019, in Advances in neural information processing systems, 8024–8035

    Paszke, A., Gross, S., Massa, F., et al. 2019, in Advances in neural information processing systems, 8024–8035

  15. [23]

    2018, Astroparticle Physics, 99, 43

    Tian, Z., Wang, Z., Liu, Y., et al. 2018, Astroparticle Physics, 99, 43

  16. [24]

    2019, in 36th International Cosmic Ray Conference (ICRC2019), Vol

    Wang, X. 2019, in 36th International Cosmic Ray Conference (ICRC2019), Vol. 36, 820

  17. [25]

    1995, Astroparticle Physics, 4, 119

    Westerhoff, S., Funk, B., Lindner, A., et al. 1995, Astroparticle Physics, 4, 119

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.