REVIEW 3 major objections 4 minor 1 cited by
Neural network biased corrections: Cautionary study in background corrections for quenched jets
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Neural-network jet corrections trained on unquenched jets are biased on quenched jets, distorting a simulated R_AA by 18-47%.
desk verdict Qualitative claim about NN background-correction bias on quenched jets is solid; the headline 18–47% RAA range is illustrative rather than robust, resting on a lightly validated brick approximation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the residual-error distribution $\delta p_{T,\mathrm{jet}} \equiv p^{\mathrm{corr}}_{T,\mathrm{jet}} - p^{\mathrm{truth}}_{T,\mathrm{jet}}$, the difference between the background-corrected jet $p_T$ and the true jet $p_T$ known from simulation. The paper tracks how the mean and width of this distribution evolve as jets are quenched in QGP bricks of increasing length, establishing that a 3.5 fm brick reproduces the substructure modification seen in the full hydrodynamically modeled events. The neural networks are trained on unquenched pp jets embedded in hydro backgrounds, and they map reco-jet parameters to truth $p_T$ using either only the area-based inputs ($p^{\mathrm{reco}}_{T,\mathrm{jet}}$, $\rho_\mathrm{bkg}$, $A_\mathrm{jet}$) or those inputs plus substructure features such as jet angularity, the number of constituents, and the $p_T$ of the leading constituents; the substructure features are what make the correction sensitive to quenching, and the sensitivity is what produces the bias.
What would settle it
Real-data embedding test: take high-$p_T$ jets of known identity, embed them into recorded central Au+Au events, apply the pp-trained network corrections exactly as in this paper, and check whether the mean residual $\delta p_{T,\mathrm{jet}}$ grows with the amount of recorded substructure modification; if the mean residual stays at zero, the claimed bias is a simulation artifact rather than a property of the correction method.
Extended reading notes
Core claim
The central claim is that neural-network background corrections trained on unquenched pp jets embedded in heavy-ion backgrounds are systematically biased when applied to quenched jets, and that the bias is not a small correction but a large, $p_T$-dependent offset. In the JETSCAPE test, the residual error $\delta p_{T,\mathrm{jet}} \equiv p^{\mathrm{corr}}_{T,\mathrm{jet}} - p^{\mathrm{truth}}_{T,\mathrm{jet}}$ shifts as brick thickness grows, with the hydrodynamically modeled events matching quenching in roughly 3.5 fm QGP bricks. When those bricks are used to build a full quenched-jet spectrum and a leading-jet $R_\mathrm{AA}$ is measured through the same unfolding procedure used experimentally, every substructure-fed network biases the result by at least 18% in every $p_T$ bin, up to about 47% for the network using constituent count; the only unbiased network is the one trained on the same parameters as the area-based method, which contains no substructure information.
Load-bearing premise
The load-bearing premise is that JETSCAPE's simulations of jet quenching—both the hydrodynamically modeled QGP and the 3.5 fm brick used for the full spectrum—faithfully capture how real quenching changes jet substructure in central Au+Au collisions at 200 GeV; if they do not, the quantified 18-47% biases would not transfer to real data.
Editorial extensions
If this is right
- Any ML background correction that uses jet substructure must assume a particular amount of quenching before it can be used to measure quenching.
- In the simulated RHIC kinematics, the area-based method returns approximately the true leading-jet $R_\mathrm{AA}$, while every substructure-based network studied is biased by 18-47%, with the largest bias coming from the network that uses the number of jet constituents.
- The bias varies with jet $p_T$, so it cannot be absorbed by a global scale factor or a single efficiency correction.
- When the amount of quenching remains ambiguous after unfolding, results should be reported as a bounded range rather than as a single value, the paper recommends.
- The paper points to two possible remedies: iterative refinement of the assumed quenching during the correction, and ML classifiers that separate fake jets from real jets without depending on substructure.
Reading between the lines
- If the mechanism is generic, other observables built from substructure-corrected jet $p_T$—dijet momentum imbalance, jet fragmentation functions, groomed jet shapes—would carry similar quenching-dependent offsets even though the paper only demonstrates the effect for $R_\mathrm{AA}$.
- A direct closure test of the explanation would retrain the same networks on jets quenched at several brick lengths; if the $R_\mathrm{AA}$ bias then disappears, the effect can be parameterized and corrected by interpolation, whereas if it persists, the mismatch is not purely due to the training sample.
- The 18-47% figures come from a simulation without detector effects or medium response, so they should be read as evidence of a large systematic risk in real measurements rather than as a prediction of the exact experimental bias.
- Because the brick-to-hydro equivalence is established using JETSCAPE's own energy-loss model, comparing the bias from an independent quenching implementation would reveal how much of the effect is generic to the logic and how much is model-specific.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies whether neural-network (NN) background corrections for jet pT, trained on unquenched proton-proton jets embedded in heavy-ion backgrounds, remain unbiased when applied to quenched jets. Using JETSCAPE simulations of central Au+Au collisions at sqrt(s_NN)=200 GeV, the authors train five NNs with different jet-substructure inputs plus an area-based baseline, evaluate the residual error distributions on quenched jets from hydrodynamic QGP events and from fixed-length QGP bricks, and then build a simulated leading-jet RAA measurement using 3.5 fm brick jets. They find that substructure-sensitive NNs produce pT-dependent biases on quenched jets, with RAA values systematically below the true quenched jet RAA, and they argue that any substructure-based background correction must presuppose an amount of quenching before quenching can be measured.
Significance. If correct, the paper identifies a real and often underappreciated risk in ML-based heavy-ion jet background subtraction: the training distribution encodes unquenched jet substructure, so applying the correction to quenched jets can bias the measured RAA. The study is useful as a cautionary benchmark for experimental analyses, especially given the upcoming RHIC run and the existing ALICE use of ML corrections at the LHC. The authors are transparent about their cuts, their training-boundary artifacts, the absence of detector response, and the leading-jet-only simplification. They also make the code and notebooks publicly available, which strengthens reproducibility. The main caveat is that the quantitative RAA bias range is derived from brick quenching, whose equivalence to hydro quenching is validated only at the level of first and second moments of the correction residual and of fragmentation functions, not at the level of the full response matrix that controls the RAA.
major comments (3)
- [Section III B 2 and IV] The quantitative RAA claim in Section V ("up to a maximum of around 47% when using NN Ncons, and no less than 18% for any pT range for any NN") is computed from jets quenched in 3.5 fm bricks. Section III A validates brick-hydro equivalence by comparing the mean and standard deviation of delta-pT,jet at selected truth-pT values (Fig. 8) and by comparing fragmentation functions (Fig. 1). The RAA result, however, is controlled by the full response matrix, including the tails of delta-pT, matching inefficiencies, and the unfolding procedure. Since the paper itself notes in the Introduction that bricks "destroy effects from variable path lengths and the evolving medium on jet quenching," the 18-47% numbers are not established as representative even of JETSCAPE hydro quenching. Please either validate the full response matrix for the 3.5 fm brick against hydro events, or explicitly present the RAA numbers as an illustration of brick quenching only and soften the unqualified summary-statement claim.
- [Table I and Section II D] Table I and its footnote state that ptruth_T,jet is "used with each NN" alongside preco_T,jet, Ajet, and rho_bkg. If ptruth were an input feature, the training would be circular and the reported nonzero biases could not arise; if, as the rest of the text indicates, ptruth is the regression target, then the table and several appendix captions must be corrected to distinguish input features from the target. Please state unambiguously that ptruth is the target and list only preco, Ajet, rho_bkg, and the substructure variables as inputs.
- [Section II D and Figures 4-5] The training-boundary artifact is acknowledged but not quantitatively separated from the quenching-induced bias. Because quenched jets shift toward the low-pT training boundary, part of the observed delta-pT bias in Figs. 7 and A.5-A.9 could reflect the learned boundary rather than substructure mismatch. The fact that NNAB shows little bias is reassuring, but a control with a wider training pT range or with training on a spectrum matched to the quenched distribution would make the central interpretation cleaner. Please add such a control or explicitly state the residual ambiguity.
minor comments (4)
- [Figure 9 vs Figure 10] The cut on the area-based corrected pT is quoted as preco_T,jet - Ajet*rho_bkg > 0 GeV/c in Figure 9 but as > 12 GeV/c in Figure 10; please make these cut definitions consistent or explain the difference.
- [Appendix A.5/A.6] The captions for Figures A.5 and A.6 are mislabeled: Figure A.5 is described as "NN: AB" while the text describes training with angularity-related inputs, and Figure A.6 is described as "NN: Ang." with a similar mismatch. Please correct the captions.
- [Appendix A.10/A.12/A.13] Several appendix captions confuse ptruth and preco, and Figures A.12/A.13 have duplicated or swapped NN labels (both subcaptions say "NNNcons" in places). Please correct these labels and repeat the input-feature list consistently.
- [Section II D] There is a typo in the first sentence of Section II D: "an set of pp jets" should read "a set of pp jets."
Circularity Check
No significant circularity: the NN bias on quenched jets is computed against independent simulation truth labels, and the RAA error is a derived output rather than a fitted input; the authors' stated limitations are model-fidelity caveats, not logical circularity.
full rationale
The derivation chain is: (1) train NNs on unquenched pp jets embedded in hydro backgrounds to map preco to ptruth; (2) apply the frozen NNs to independently generated quenched jets (hydro and brick) and measure δpT,jet ≡ pcorr − ptruth against the simulation truth labels; (3) calibrate a 3.5 fm brick proxy to hydro using the mean and standard deviation of δpT,jet and the fragmentation function (Fig. 1, Fig. 8); (4) generate a full brick-quenched spectrum and propagate the measured δpT,jet bias through an unfolding-based simulated RAA measurement. The central claim — NN corrections trained on unquenched jets are biased on quenched jets, up to 47% in simulated RAA — is an output of this chain. Nothing in the claim is defined in terms of the target: the NNs never see quenched jets in training, the truth RAA is taken directly from brick simulation truth (independent of NN outputs), and the 3.5 fm calibration is a diagnostic choice (matching ⟨δpT,jet⟩ and σ(δpT,jet)), not the predicted quantity (the fractional RAA bias). The RAA bias could in principle differ from the moment-based calibration signal; it is computed, not assumed. The only self-citation is the JETSCAPE framework [26] (a co-author paper, with tunes from [27,28]); this is a public simulation code used as a tool whose output the analysis then computes on, not a conclusion imported from the citation, so it is not load-bearing in the circularity sense. The paper's own stated limitations — no jet-medium response modeled (Sec. I; 'They do not, however, model medium response to the jets'; Sec. V: 'jet-medium interactions are not captured in the simulations used'), fixed brick path length destroying variable path-length and evolving-medium effects (Sec. I: 'it also destroys effects from variable path lengths and the evolving medium on jet quenching'), and the brick-to-hydro equivalence validated on first moments of δpT,jet rather than the full response matrix that governs the RAA — are model-fidelity and transferability concerns, as the skeptic headline notes, not instances where a prediction reduces to its input by construction. Score 1 reflects one minor co-author self-citation and a simulation-based validity caveat, with no step of the derivation equivalent to its own input.
Assumptions & free parameters
free parameters (6)
- NN architecture and training protocol =
3 dense layers (100, 50, 50), ReLU, 12 epochs
- Training pT range =
0-60 GeV/c flat spectrum
- QGP brick length =
3.5 fm
- Background density estimator cuts =
Two highest pT jets removed for rho; pcorr > 0 GeV/c cut
- Leading jet selection and pT thresholds =
Leading jet, |eta| < 1, truth pT > 12 GeV/c, reco-area rho > 0 (or >12 in Fig. 10)
- Unfolding iterations =
4 iterations Bayesian, or 1-bin efficiency
assumptions (6)
- standard math Anti-kT jet clustering is infrared and collinear safe and yields stable areas for R=0.4 jets
- domain assumption JETSCAPE with the stated tune provides a realistic simulation of pp and Au+Au events at sqrt(s)=200 GeV
- domain assumption The hydrodynamically modeled QGP events (3,100 Au+Au, hadronized 10x) produce realistic background particle distributions
- ad hoc to paper Quenching in a 3.5 fm static QGP brick is equivalent, for substructure modification, to quenching in hydro events
- domain assumption The experimental practice of constructing the response matrix from unquenched pp MC, rather than quenched MC, is the correct comparator
- domain assumption No detector effects or efficiencies are needed to estimate the NN-induced bias
Cite this review
Pith. "Pith review of Neural network biased corrections: Cautionary study in background corrections for quenched jets." pith.science (2026). https://pith.science/paper/OJG5D74C
@misc{pith2026241215440,
author = {Pith},
title = {Pith review of: Neural network biased corrections: Cautionary study in background corrections for quenched jets},
year = {2026},
howpublished = {\url{https://pith.science/paper/OJG5D74C}},
note = {Machine review of arXiv:2412.15440}
}
abstract
Jets clustered from heavy ion collision measurements combine a dense background of particles with those actually resulting from a hard partonic scattering. The background contribution to jet transverse momentum ($p_{T}$) may be corrected by subtracting the collision average background; however, the background inhomogeneity limits the resolution of this correction. Many recent studies have embedded jets into heavy ion backgrounds and demonstrated a markedly improved background correction is achievable by using neural networks (NNs) trained with aspects of jet substructure which are used to map measured jet $p_\mathrm{T}$ to the embedded truth jet $p_\mathrm{T}$. However, jet quenching in heavy ion collisions modifies jet substructure, and correspondingly biases the NNs' background corrections. This study investigates those biases by using simulations of jet quenching in central Au+Au collisions at $\sqrt{s_\mathrm{NN}}=200\;\mathrm{GeV}/c$ with hydrodynamically modeled quark-gluon plasma (QGP) evolution. To demonstrate the magnitude of the effect of such biases in measurement, a leading jet nuclear modification factor ($R_\mathrm{AA}$) is calculated and reported using the NN background correction on jets quenched utilizing a brick of QGP.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 1 Pith paper
-
High-Dimensional Unfolding in Large Backgrounds
OmniFold-HI, an ML unfolding algorithm that handles large backgrounds and high-dimensional auxiliary observables, is derived, shown equivalent to iterative Bayesian unfolding, and demonstrated to improve jet-substruct...
Reference graph
Works this paper leans on
-
[1]
Use only the highest- pT IP scattering for all hard scatterings in each event
-
[2]
Cut events with IP pseudorapidity ηIP > |1.0|
-
[3]
Cluster all final-state particles resulting from the selected IP into anti- kT jets. Consider all jets relative to the IP within distance ∆R ≡ p (ηjet − ηIP)2 + (ϕjet − ϕIP)2 of ∆ R < 0.4. Discard the event if there are no such jets. If there are, select the highest- pT of these jet as the truth jet with ptruth T,jet
-
[4]
In hydro events, use the background particles from the same hydro event in the clustering
Cluster the background particles into kT jets [29]. In hydro events, use the background particles from the same hydro event in the clustering. In all other events, use the background particles saved from one of the hydro events
-
[5]
Remove the two highest pT jets, and record the median jet-pT density (pT,jet/Ajet) as ρbkg
-
[6]
Cluster the constituents from the IP scatterings and the background particles together into anti- kT jets. Select all resulting jets that are within ∆ R < 0.3 of the truth jet. If there are none, discard the event. Otherwise, the highest-pT of these jets is the “reco jet” with preco T,jet
-
[7]
Record ptruth T,jet , preco T,jet, and other event parameters used to train Neural Networks. A list of which pa- rameters are used to train each neural network is given in Table I. TABLE I. Neural network (NN) training parameters Label Additional Training Parameters † NNAB (none) NNAng Angularity: α ≡ P i pT,i∆Ri, where i runs over all constituents, and ∆...
-
[8]
Algorithm to Measure Jets Quenching in Experiments a. Measured data consists of events with jets con- stituents clustered together with the heavy back- ground resulting from heavy ion collisions. This clustering results in detector-level jets, with a spec- trum of preco T,jet (i.e. d preco T,jet/dpT). b. The reco-jets are background corrected, commonly us...
Show all 45 references
-
[9]
measurement data
Algorithm Used in this Paper to Measure Jets The algorithm implemented for the results reported in this paper are comparable to that listed in Sec. III B 1. The differences listed below. a. The “measurement data” consists of JETSCAPE simulated jets quenched in 3 .5 fm bricks o...
-
[10]
Hardware All code was run on a single machine equipped with an AMD Ryzen Threadripper 3960X Processor, two NVIDIA GeForce RTX 3090 GPUs, and 128 GB of DDR4 ram
-
[11]
Input files, scripts, and codes, are archived at github.site
Software The following software process was used. Input files, scripts, and codes, are archived at github.site. • JETSCAPE 3.6.4 [26] was pulled from online at https://github.com/JETSCAPE/JETSCAPE, compiled locally, and run with XML files. • JETSCAPE’s output .dat.gz files wer...
-
[12]
J. W. Harris and B. Muller, Ann. Rev. Nucl. Part. Sci. 46, 71 (1996), arXiv:hep-ph/9602235
1996 arXiv
-
[13]
Arsene et al
I. Arsene et al. (BRAHMS), Nucl. Phys. A 757, 1 (2005), arXiv:nucl-ex/0410020
2005 arXiv
-
[14]
B. B. Back et al. (PHOBOS), Nucl. Phys. A 757, 28 (2005), arXiv:nucl-ex/0410022
2005 arXiv
- [15]
-
[16]
Adcox et al
K. Adcox et al. (PHENIX), Nucl. Phys. A 757, 184 (2005), arXiv:nucl-ex/0410003
2005 arXiv
-
[17]
Aad et al
G. Aad et al. (ATLAS), Phys. Rev. Lett. 105, 252303 (2010), arXiv:1011.6182 [hep-ex]
2010 arXiv
-
[18]
Chatrchyan et al
S. Chatrchyan et al. (CMS), Phys. Rev. C 84, 024906 (2011), arXiv:1102.1957 [nucl-ex]
2011 arXiv
-
[19]
Aamodt et al
K. Aamodt et al. (ALICE), Phys. Rev. Lett. 105, 252302 (2010), arXiv:1011.3914 [nucl-ex]
2010 arXiv
- [20]
- [21]
-
[22]
Adams et al
J. Adams et al. (STAR), Phys. Rev. Lett. 93, 252301 (2004), arXiv:nucl-ex/0407007
2004 arXiv
-
[23]
Brock et al
R. Brock et al. (CTEQ), Rev. Mod. Phys. 67, 157 (1995)
1995
-
[24]
Cunqueiro and A
L. Cunqueiro and A. M. Sickles, Prog. Part. Nucl. Phys. 124, 103940 (2022), arXiv:2110.14490 [nucl-ex]
2022 arXiv
-
[25]
N. J. Abdulameer et al. (PHENIX), (2024), arXiv:2408.11144 [hep-ex]
2024
-
[26]
Abdulhamid et al
M. Abdulhamid et al. (STAR), Phys. Rev. C 109, 044909 (2024), arXiv:2307.13891 [nucl-ex]
2024
-
[27]
Abdulhamid et al
M. Abdulhamid et al. (STAR), Phys. Rev. C 110, 044908 (2024), arXiv:2404.08784 [nucl-ex]
2024
- [28]
-
[29]
Cacciari and G
M. Cacciari and G. P. Salam, Phys. Lett. B 659, 119 (2008), arXiv:0707.1378 [hep-ph]
2008 arXiv
-
[30]
Cacciari, G
M. Cacciari, G. P. Salam, and G. Soyez, Eur. Phys. J. C 72, 1896 (2012), arXiv:1111.6097 [hep-ph]
2012 arXiv
-
[31]
Sj¨ ostrand, Computer Physics Communications246, 106910 (2020)
T. Sj¨ ostrand, Computer Physics Communications246, 106910 (2020)
2020
-
[32]
Bierlich et al., SciPost Phys
C. Bierlich et al., SciPost Phys. Codeb. 2022, 8 (2022), arXiv:2203.11601 [hep-ph]
2022 arXiv
-
[33]
Bellm et al., Eur
J. Bellm et al., Eur. Phys. J. C 76, 196 (2016), arXiv:1512.01178 [hep-ph]
2016 arXiv
-
[34]
Haake and C
R. Haake and C. Loizides, Phys. Rev. C 99, 064904 (2019), arXiv:1810.06324 [nucl-ex]
2019 arXiv
-
[35]
Acharya et al
S. Acharya et al. (ALICE), Phys. Lett. B 849, 138412 (2024), arXiv:2303.00592 [nucl-ex]
2024 arXiv
-
[36]
Mengel, P
T. Mengel, P. Steffanic, C. Hughes, A. C. O. Da Silva, and C. Nattrass, (2024), arXiv:2402.10945 [hep-ex]
2024 arXiv
-
[37]
J. H. Putschke et al., (2019), arXiv:1903.07706 [nucl-th]
2019 arXiv
-
[38]
Everett et al
D. Everett et al. (JETSCAPE), Phys. Rev. C 103, 054904 (2021), arXiv:2011.01430 [hep-ph]
2021 arXiv
-
[39]
Kumar et al
A. Kumar et al. (JETSCAPE), Phys. Rev. C 107, 034911 (2023), arXiv:2204.01163 [hep-ph]
2023 arXiv
-
[40]
S. D. Ellis and D. E. Soper, Phys. Rev. D 48, 3160 (1993), arXiv:hep-ph/9305266
1993 arXiv
-
[41]
Abadi et al., (2016), arXiv:1603.04467 [cs.DC]
M. Abadi et al., (2016), arXiv:1603.04467 [cs.DC]
2016 arXiv
-
[42]
Agostinelli et al
S. Agostinelli et al. (GEANT4), Nucl. Instrum. Meth. A 506, 250 (2003)
2003
-
[43]
Adye, in PHYSTAT 2011 (CERN, Geneva, 2011) pp
T. Adye, in PHYSTAT 2011 (CERN, Geneva, 2011) pp. 313–318, arXiv:1105.1160 [physics.data-an]
2011 arXiv
-
[44]
R. Brun, F. Rademakers, and S. Panacek, in CERN School of Computing (CSC 2000)(2000) pp. 11–42
2000
-
[45]
Pedregosa et al., J
F. Pedregosa et al., J. Machine Learning Res. 12, 2825 (2011), arXiv:1201.0490 [cs.LG]
2011 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.