{"id":"48ffad75-fa0e-485e-a4b0-139a963d8019","arxiv_id":"2412.15440","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Neural network jet background corrections trained on unquenched jets are systematically biased for quenched jets, causing simulated RAA measurements to be up to 47% too low.","lead":"A simulation study shows that neural network corrections, which improve jet background subtraction in heavy ion collisions when trained on ordinary proton-proton collisions, become biased when applied to quenched jets in quark-gluon plasma. The paper quantifies this bias and warns that it can distort the jet quenching measurement (RAA) by up to 47%.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 18–47% RAA bias range is anchored to 3.5 fm brick jets, but the brick-to-hydro equivalence is validated only by first moments of δpT, not by the full response matrix that controls the RAA.","rationale":"The paper's qualitative central claim—that NN background corrections trained on unquenched jets are biased when applied to quenched jets—is logically sound and would survive even if the brick/hydro match failed, because it follows from any modification of the substructure features used by the network. The load-bearing part is the quantitative RAA error range, which is the paper's stated demonstration of 'the magnitude of the effect.' That magnitude is generated with the brick proxy, so the validity of the brick-hydro equivalence is the key unverified link. The current validation, using only the mean and standard deviation of δpT at a few truth-pT values, is insufficient for the full response matrix and unfolding procedure that determine the RAA. This concern is partly internal: even within JETSCAPE, the fixed-length 3.5 fm brick may not reproduce the hydro quenched-jet response matrix, regardless of whether JETSCAPE matches real data. The reader's weakest assumption about JETSCAPE fidelity is related but broader; our concern is more specific and testable. Because the quantitative headline is conditional on this equivalence, the existing CONDITIONAL verdict is appropriate and no further change is needed.","tokens_in":17702,"tokens_out":6490,"duration_ms":60816,"concrete_test":"Generate RLeadJet_AA from the 31,000 hydro-quenched jets directly, embedding them in the same hydro backgrounds and applying the same five NN corrections and unfolding pipeline used for the brick sample; compare the resulting per-pT RAA to the 3.5 fm brick curves in Figs. 9–10. If the hydro-based bias falls outside the quoted 18–47% range or deviates by more than about 10 points in any pT bin, the brick proxy is not sufficiently validated for the headline numbers.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—'up to around 47% ... no less than 18% for any pT range for any NN' (Sec. V)—rests on the RAA simulation of Sec. III B 2, which uses jets quenched in fixed 3.5 fm QGP bricks embedded in hydro backgrounds. The brick-to-hydro equivalence is validated in Sec. III A only by comparing the mean and standard deviation of the correction residual δpT,jet for a few truth-pT values (Fig. 8), plus the fragmentation comparison in Fig. 1. The RAA result, however, is controlled by the full response matrix—including tails, pT bins below 20 GeV/c, matching inefficiencies, and unfolding—and a brick with fixed path length cannot be assumed to reproduce those jointly. The authors themselves note that bricks 'destroy effects from variable path lengths and the evolving medium on jet quenching' (Sec. III). Consequently, the 18–47% range is not yet established as representative even of JETSCAPE hydro quenching, independent of the separate question of whether JETSCAPE matches real Au+Au data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies whether neural-network (NN) background corrections for jet pT, trained on unquenched proton-proton jets embedded in heavy-ion backgrounds, remain unbiased when applied to quenched jets. Using JETSCAPE simulations of central Au+Au collisions at sqrt(s_NN)=200 GeV, the authors train five NNs with different jet-substructure inputs plus an area-based baseline, evaluate the residual error distributions on quenched jets from hydrodynamic QGP events and from fixed-length QGP bricks, and then build a simulated leading-jet RAA measurement using 3.5 fm brick jets. They find that substructure-sensitive NNs produce pT-dependent biases on quenched jets, with RAA values systematically below the true quenched jet RAA, and they argue that any substructure-based background correction must presuppose an amount of quenching before quenching can be measured.","tokens_in":17911,"tokens_out":6486,"duration_ms":62792,"significance":"If correct, the paper identifies a real and often underappreciated risk in ML-based heavy-ion jet background subtraction: the training distribution encodes unquenched jet substructure, so applying the correction to quenched jets can bias the measured RAA. The study is useful as a cautionary benchmark for experimental analyses, especially given the upcoming RHIC run and the existing ALICE use of ML corrections at the LHC. The authors are transparent about their cuts, their training-boundary artifacts, the absence of detector response, and the leading-jet-only simplification. They also make the code and notebooks publicly available, which strengthens reproducibility. The main caveat is that the quantitative RAA bias range is derived from brick quenching, whose equivalence to hydro quenching is validated only at the level of first and second moments of the correction residual and of fragmentation functions, not at the level of the full response matrix that controls the RAA.","major_comments":[{"comment":"The quantitative RAA claim in Section V (\"up to a maximum of around 47% when using NN Ncons, and no less than 18% for any pT range for any NN\") is computed from jets quenched in 3.5 fm bricks. Section III A validates brick-hydro equivalence by comparing the mean and standard deviation of delta-pT,jet at selected truth-pT values (Fig. 8) and by comparing fragmentation functions (Fig. 1). The RAA result, however, is controlled by the full response matrix, including the tails of delta-pT, matching inefficiencies, and the unfolding procedure. Since the paper itself notes in the Introduction that bricks \"destroy effects from variable path lengths and the evolving medium on jet quenching,\" the 18-47% numbers are not established as representative even of JETSCAPE hydro quenching. Please either validate the full response matrix for the 3.5 fm brick against hydro events, or explicitly present the RAA numbers as an illustration of brick quenching only and soften the unqualified summary-statement claim.","section":"Section III B 2 and IV"},{"comment":"Table I and its footnote state that ptruth_T,jet is \"used with each NN\" alongside preco_T,jet, Ajet, and rho_bkg. If ptruth were an input feature, the training would be circular and the reported nonzero biases could not arise; if, as the rest of the text indicates, ptruth is the regression target, then the table and several appendix captions must be corrected to distinguish input features from the target. Please state unambiguously that ptruth is the target and list only preco, Ajet, rho_bkg, and the substructure variables as inputs.","section":"Table I and Section II D"},{"comment":"The training-boundary artifact is acknowledged but not quantitatively separated from the quenching-induced bias. Because quenched jets shift toward the low-pT training boundary, part of the observed delta-pT bias in Figs. 7 and A.5-A.9 could reflect the learned boundary rather than substructure mismatch. The fact that NNAB shows little bias is reassuring, but a control with a wider training pT range or with training on a spectrum matched to the quenched distribution would make the central interpretation cleaner. Please add such a control or explicitly state the residual ambiguity.","section":"Section II D and Figures 4-5"}],"minor_comments":[{"comment":"The cut on the area-based corrected pT is quoted as preco_T,jet - Ajet*rho_bkg > 0 GeV/c in Figure 9 but as > 12 GeV/c in Figure 10; please make these cut definitions consistent or explain the difference.","section":"Figure 9 vs Figure 10"},{"comment":"The captions for Figures A.5 and A.6 are mislabeled: Figure A.5 is described as \"NN: AB\" while the text describes training with angularity-related inputs, and Figure A.6 is described as \"NN: Ang.\" with a similar mismatch. Please correct the captions.","section":"Appendix A.5/A.6"},{"comment":"Several appendix captions confuse ptruth and preco, and Figures A.12/A.13 have duplicated or swapped NN labels (both subcaptions say \"NNNcons\" in places). Please correct these labels and repeat the input-feature list consistently.","section":"Appendix A.10/A.12/A.13"},{"comment":"There is a typo in the first sentence of Section II D: \"an set of pp jets\" should read \"a set of pp jets.\"","section":"Section II D"}],"recommendation":"major_revision","confidential_remarks":"The paper is a suitable cautionary study for physics.data-an and would likely be of interest to the heavy-ion jet community. The central qualitative point is sound and the authors are appropriately transparent about many limitations. My main concern is that the most prominent quantitative claim (the 18-47% RAA bias) relies on a brick-to-hydro equivalence that has only been checked at the level of residual means and widths, not at the level of the response matrix that actually controls the unfolded RAA. This is fixable by either adding a hydro-based RAA validation or rewording the claims, so I recommend major revision rather than rejection. The Table I ambiguity about ptruth as an input versus target should also be resolved before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is a serious, honest cautionary study, and the qualitative claim is solid. The headline quantitative claim—RAA bias of 18–47%—is real, but it is anchored to a brick approximation that is validated only at the level of δpT means and widths, not the full response matrix. Treat the number as illustrative of the mechanism, not as a robust prediction for RHIC.\n\nWhat is actually new: prior work (Haake-Loizides, ALICE, Mengel et al.) established NN background corrections, but this is the first systematic look at what happens when those corrections are applied to quenched jets. The authors train on unquenched pp jets and then apply to JETSCAPE jets quenched in hydrodynamically modeled QGP backgrounds, which is a meaningful step up from toy backgrounds. The result—that the correction bias grows with quenching and depends on the substructure variable—is well supported by the residual distributions and the RAA comparisons. The paper is also transparent about the leading-jet-only selection, the training-boundary artifacts, and the absence of detector response. That transparency earns trust.\n\nSoft spots, in proportion. First, the 18–47% range comes from jets quenched in fixed 3.5 fm bricks embedded in hydro backgrounds. The authors themselves note that bricks kill variable path lengths and the evolving medium. The brick-to-hydro equivalence is checked only via the mean and standard deviation of δpT (Fig. 8) and a fragmentation comparison. But the RAA result depends on the full response matrix, including tails, matching inefficiencies, and unfolding. So the quantitative claim is not yet established even within JETSCAPE. That does not hurt the qualitative conclusion, but the abstract should not let readers think the number is firmer than it is. Second, Table I's footnote lists ptruth as a training parameter; in context it is clearly the regression target, but the wording invites a leakage accusation and should be fixed. Third, the code link is malformed and no data or trained weights are released; for a simulation-only cautionary study, that is a moderate reproducibility gap. Fourth, only the leading jet per event is used, which the authors say likely underestimates the effect.\n\nVerdict: deserves a serious referee. The core message will survive review; the quantitative claim needs qualification or a fuller validation of the brick-to-hydro mapping. I would take it to a reading group and would cite it if I work on jet background corrections.\n\nRecommendation: send to peer review, with a request that the authors either soften the RAA claim or validate the brick equivalence at the response-matrix level.","headline":"Qualitative claim about NN background-correction bias on quenched jets is solid; the headline 18–47% RAA range is illustrative rather than robust, resting on a lightly validated brick approximation.","tokens_in":18475,"tokens_out":2846,"would_cite":true,"duration_ms":25033,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["25.75.-q","07.05.Mh"],"model":"deepseek-v4-flash","headline":"Neural-network jet corrections trained on unquenched jets are biased on quenched jets, distorting a simulated R_AA by 18-47%.","keywords":["neural networks","jet quenching","background subtraction","jet substructure","heavy-ion collisions","nuclear modification factor","RAA","JETSCAPE"],"falsifier":"Real-data embedding test: take high-$p_T$ jets of known identity, embed them into recorded central Au+Au events, apply the pp-trained network corrections exactly as in this paper, and check whether the mean residual $\\delta p_{T,\\mathrm{jet}}$ grows with the amount of recorded substructure modification; if the mean residual stays at zero, the claimed bias is a simulation artifact rather than a property of the correction method.","tokens_in":17435,"feed_emoji":"⚛️","tokens_out":10868,"duration_ms":82734,"temperature":0.7,"pith_summary":"Jets in heavy-ion collisions are measured on top of a huge background of soft particles, and neural networks that use jet substructure have been shown to correct for that background more precisely than simple area-based subtraction. This paper argues that such networks carry a hidden assumption: because they are trained on unquenched proton-proton jets, applying them to jets quenched in the quark-gluon plasma biases the correction, since quenching changes exactly the substructure variables the network relies on. Using JETSCAPE simulations of central Au+Au collisions at $\\sqrt{s_{NN}}=200$ GeV and faster \"brick\" simulations of quenching, the bias is shown to grow with quenching and to propagate into a mock leading-jet $R_\\mathrm{AA}$ measurement, lowering the extracted ratio by 18-47% depending on the substructure inputs. The paper's conclusion is a caution: any substructure-based ML background correction presupposes an amount of quenching, so ambiguous results should be reported as bounded ranges rather than single values.","feed_headline":"ML jet corrections skew quenched-jet measurements by up to 47%","feed_subtitle":"Networks trained on unquenched jets assume the quenching they are meant to measure.","key_machinery":"The load-bearing object is the residual-error distribution $\\delta p_{T,\\mathrm{jet}} \\equiv p^{\\mathrm{corr}}_{T,\\mathrm{jet}} - p^{\\mathrm{truth}}_{T,\\mathrm{jet}}$, the difference between the background-corrected jet $p_T$ and the true jet $p_T$ known from simulation. The paper tracks how the mean and width of this distribution evolve as jets are quenched in QGP bricks of increasing length, establishing that a 3.5 fm brick reproduces the substructure modification seen in the full hydrodynamically modeled events. The neural networks are trained on unquenched pp jets embedded in hydro backgrounds, and they map reco-jet parameters to truth $p_T$ using either only the area-based inputs ($p^{\\mathrm{reco}}_{T,\\mathrm{jet}}$, $\\rho_\\mathrm{bkg}$, $A_\\mathrm{jet}$) or those inputs plus substructure features such as jet angularity, the number of constituents, and the $p_T$ of the leading constituents; the substructure features are what make the correction sensitive to quenching, and the sensitivity is what produces the bias.","core_discovery":"The central claim is that neural-network background corrections trained on unquenched pp jets embedded in heavy-ion backgrounds are systematically biased when applied to quenched jets, and that the bias is not a small correction but a large, $p_T$-dependent offset. In the JETSCAPE test, the residual error $\\delta p_{T,\\mathrm{jet}} \\equiv p^{\\mathrm{corr}}_{T,\\mathrm{jet}} - p^{\\mathrm{truth}}_{T,\\mathrm{jet}}$ shifts as brick thickness grows, with the hydrodynamically modeled events matching quenching in roughly 3.5 fm QGP bricks. When those bricks are used to build a full quenched-jet spectrum and a leading-jet $R_\\mathrm{AA}$ is measured through the same unfolding procedure used experimentally, every substructure-fed network biases the result by at least 18% in every $p_T$ bin, up to about 47% for the network using constituent count; the only unbiased network is the one trained on the same parameters as the area-based method, which contains no substructure information.","pith_inferences":["If the mechanism is generic, other observables built from substructure-corrected jet $p_T$—dijet momentum imbalance, jet fragmentation functions, groomed jet shapes—would carry similar quenching-dependent offsets even though the paper only demonstrates the effect for $R_\\mathrm{AA}$.","A direct closure test of the explanation would retrain the same networks on jets quenched at several brick lengths; if the $R_\\mathrm{AA}$ bias then disappears, the effect can be parameterized and corrected by interpolation, whereas if it persists, the mismatch is not purely due to the training sample.","The 18-47% figures come from a simulation without detector effects or medium response, so they should be read as evidence of a large systematic risk in real measurements rather than as a prediction of the exact experimental bias.","Because the brick-to-hydro equivalence is established using JETSCAPE's own energy-loss model, comparing the bias from an independent quenching implementation would reveal how much of the effect is generic to the logic and how much is model-specific."],"forward_implications":["Any ML background correction that uses jet substructure must assume a particular amount of quenching before it can be used to measure quenching.","In the simulated RHIC kinematics, the area-based method returns approximately the true leading-jet $R_\\mathrm{AA}$, while every substructure-based network studied is biased by 18-47%, with the largest bias coming from the network that uses the number of jet constituents.","The bias varies with jet $p_T$, so it cannot be absorbed by a global scale factor or a single efficiency correction.","When the amount of quenching remains ambiguous after unfolding, results should be reported as a bounded range rather than as a single value, the paper recommends.","The paper points to two possible remedies: iterative refinement of the assumed quenching during the correction, and ML classifiers that separate fake jets from real jets without depending on substructure."],"supporting_citations":[{"why":"Prior work showing NNs using jet substructure improve embedded-jet pT corrections; the baseline this paper reconfirms for unquenched jets.","marker":"[23]"},{"why":"A related embedding study extending the NN substructure-correction approach; one of the methods this paper adapts.","marker":"[25]"},{"why":"The JETSCAPE framework that generated the pp, hydro, and brick jet samples.","marker":"[26]"},{"why":"The JETSCAPE tune parameters used for the 200 GeV pp and Au+Au simulations.","marker":"[27, 28]"},{"why":"The area-based background subtraction method used as the baseline correction.","marker":"[18]"},{"why":"FastJet, used for clustering, rho_bkg measurement, and jet matching.","marker":"[19]"},{"why":"The LHC R_AA measurement using ML-based corrections, the experimental context the authors compare against.","marker":"[24]"},{"why":"RooUnfold, used for the Bayesian unfolding in the mock R_AA measurement.","marker":"[32]"}],"fun_headline_variants":["Neural net jet corrections bias quenched-jet R_AA by 18-47%","ML jet corrections skewed by jet quenching, up to 47% in R_AA","Quenched jet substructure biases neural network background corrections","Unquenched-trained NNs mis-correct quenched jets by up to 47%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that JETSCAPE's simulations of jet quenching—both the hydrodynamically modeled QGP and the 3.5 fm brick used for the full spectrum—faithfully capture how real quenching changes jet substructure in central Au+Au collisions at 200 GeV; if they do not, the quantified 18-47% biases would not transfer to real data.","fun_headline_variants_meta":{"raw":{"variants":["Neural net jet corrections bias quenched-jet R_AA by 18-47%","ML jet corrections skewed by jet quenching, up to 47% in R_AA","Quenched jet substructure biases neural network background corrections","Unquenched-trained NNs mis-correct quenched jets by up to 47%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000888,"raw_usage":{"total_tokens":3847,"prompt_tokens":978,"completion_tokens":2869,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":594,"completion_tokens_details":{"reasoning_tokens":2791}},"tokens_in":594,"tokens_out":2869,"duration_ms":21718,"temperature":1.0,"reasoning_tokens":2791,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:26:25.121139+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Real-data embedding test: take high-$p_T$ jets of known identity, embed them into recorded central Au+Au events, apply the pp-trained network corrections exactly as in this paper, and check whether the mean residual $\\delta p_{T,\\mathrm{jet}}$ grows with the amount of recorded substructure modification; if the mean residual stays at zero, the claimed bias is a simulation artifact rather than a property of the correction method.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A related embedding study extending the NN substructure-correction approach; one of the methods this paper adapts."},{"cited_title":"Abdulhamid et al","cited_arxiv_id":null,"evidence_quote":"The JETSCAPE framework that generated the pp, hydro, and brick jet samples."}],"review_version":1}