{"id":"8623454c-7aba-4d1b-9ff4-7080ff0bf42f","arxiv_id":"1908.05196","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A particle-based deep neural network with standardized inputs and principal component analysis is projected to improve the expected significance for longitudinally polarized ZZ scattering at the HL-LHC to about 1.7 standard deviations.","lead":"Using simulated LHC collisions, this paper shows that a particular deep neural network can pick out rare collisions where two Z bosons scatter with both particles polarized lengthwise, a process tied to the Higgs mechanism. It projects that with the full High-Luminosity LHC dataset this signal could reach about 1.7 standard deviations, a modest but real improvement over previous methods.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pileup neglect is the load-bearing gap: the 1.74σ HL-LHC projection is a no-pileup estimate, and the DNN's jet and isolation inputs are exactly what pileup distorts.","rationale":"The reader's weakest assumption identifies the same gap: pileup is neglected, and the central projection is therefore conditional on a no-pileup environment. This is the most load-bearing concern because it is explicit, it directly affects the inputs used by the DNN and the event selection, and the claimed improvement is small enough to be erased by realistic pileup. The paper's internal comparisons are otherwise coherent: the BDT baseline of 1.41σ reproduces the prior CMS projection of about 1.4σ, and the DNN-PCA significances follow the described pipeline. The other issues noted by the reader, such as the flat 10% systematic and the absence of a trials correction for the best-performing configuration, are real but secondary; they would change the headline number at the few-tenths-of-a-sigma level rather than invalidate the method. A pileup rerun is a concrete, decisive check: if the DNN-PCA significance under pileup remains above the BDT baseline, the central claim is supported; if it drops below, the claim is not supported. The verdict should remain CONDITIONAL pending that check, so no adjustment to the reader's verdict is needed.","tokens_in":7989,"tokens_out":3601,"duration_ms":41443,"concrete_test":"Re-run the identical MadGraph/Pythia/Delphes pipeline with CMS HL-LHC pileup, for example 200 inelastic collisions per bunch crossing as in the CMS HL-LHC Delphes card, while keeping the same event selection, particle-based DNN, YJ&STD preprocessing, PCA rotation, and three-dimensional PC fit. Report the Asimov significance for DNN-PCA and for the BDT baseline under the same pileup conditions. If the DNN-PCA significance falls below the BDT baseline or below roughly 1.4σ, the paper's central claim does not survive; also report shifts in jet multiplicity, mjj, and b-veto efficiency to identify the physical mechanism.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, 1.74σ expected significance for the longitudinal ZZ fraction at 3000 fb−1, depends on the DNN-PCA discrimination surviving realistic HL-LHC conditions. The paper states that Delphes was configured to simulate the CMS HL-LHC detector, but 'pileup has been neglected as in ref. [24]'. This is a load-bearing omission because the event selection and the DNN inputs are dominated by pileup-sensitive objects: at least two jets with pT > 25 GeV and |η| < 4.7, mjj > 400 GeV, |Δηjj| > 2.4, no b-tagged jets, and four isolated leptons. With 140–200 pileup interactions per bunch crossing, additional jets can alter the selected jet pair or fail the b-veto; isolation and lepton pT can be degraded. Since the DNN consumes the four-momenta of exactly these leptons and jets, the learned LL/TL/TT/qqZZ/ggZZ scores and the PC templates used in the multidimensional fit are likely to shift. Therefore the quoted 1.74σ is a no-pileup projection, not a direct HL-LHC expectation. Because the claimed gain over the BDT baseline (1.41σ → 1.74σ) is modest, even a moderate pileup-induced degradation could erase the claimed improvement.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the extraction of the longitudinal (LL) polarization fraction in vector-boson-scattering ZZ -> 4l events at the HL-LHC using fast simulation. The authors generate LL, TL, TT, qqZZ, and ggZZ samples with MadGraph/Pythia/Delphes, apply a VBS-like event selection, and compare a BDT, a dense DNN, and a particle-based DNN with preprocessing by standardization and Yeo-Johnson transformation, followed by principal-component analysis of the DNN outputs. They report expected significances from the Asimov formula, with the best result being 1.74 sigma (1.66 sigma with a flat 10% systematic uncertainty) at 3000 fb^-1 using a three-dimensional fit to PC1-PC3, compared with 1.41 sigma for their BDT baseline. The paper concludes that the method improves the LL sensitivity and is applicable to other helicity-fraction measurements.","tokens_in":8233,"tokens_out":6408,"duration_ms":65137,"significance":"If the projection is robust, the paper offers a useful, relatively simple improvement over the previous CMS BDT projection of about 1.4 sigma for a challenging measurement, and the comparison across preprocessing and model variants is informative. Its strengths are that all significance values are listed in Table I and Fig. 8, the Asimov significance procedure is standard, and the BDT baseline is explicitly validated against the CMS result. The main weaknesses are that the simulation neglects pileup, which directly affects the inputs on which the event selection and the DNN rely, and that the systematic uncertainty is an unvalidated flat 10% assumption; both make the quantitative HL-LHC projection, rather than the qualitative methodological conclusion, uncertain.","major_comments":[{"comment":"The simulation paragraph explicitly states that 'pileup' has been neglected as in ref. [24]. This is load-bearing for the central claim because the event selection requires at least two jets with pT > 25 GeV, |eta| < 4.7, mjj > 400 GeV, |Delta eta_jj| > 2.4, no b-tagged jets, and four isolated leptons, and the DNN consumes exactly these lepton and jet four-momenta. At the HL-LHC with 140-200 pileup interactions per bunch crossing, pileup can add jets, alter the selected jet pair, break the b-veto, and degrade lepton isolation, shifting the DNN scores and the PC templates used in the 1.74 sigma fit. Because the improvement over the BDT baseline (1.41 sigma to 1.74 sigma) is modest, even a moderate pileup-induced degradation could erase the claimed gain. Please rerun with pileup overlay, or at least with a robustness test against isolation and jet-threshold variations, or clearly restate the result as a no-pileup feasibility projection.","section":"Simulation and event selection paragraph"},{"comment":"The 'Stat. & syst.' row applies a flat 10% uncertainty to both signal and background yields, but the paper gives no derivation or validation for this value. The quoted significance with systematics is therefore not a measured property of the method but an input assumption; different correlations or uncertainties for qqZZ, ggZZ, and signal would materially change the 1.66 sigma number. The authors should justify the 10% value, test its variation (for example 5% and 20%), or present the result as statistical-only with the systematic treated as an illustrative prescription.","section":"Table I and preceding text"},{"comment":"The paper reports 1.65 sigma for a PC1-PC2 fit and 1.74 sigma for a PC1-PC3 fit but does not define the test statistic, binning, or likelihood used for the Asimov calculation on multidimensional histograms. Without this information, a reader cannot reproduce the central number or assess whether the gain over using PC1 alone (1.55 sigma) is meaningful. Please specify the exact fit procedure (for example, binned Poisson likelihood, number of bins per dimension, and nuisance parametrization) and, if feasible, show the sensitivity of the significance to the binning choice.","section":"PC fit paragraph and Fig. 7"},{"comment":"The final significance is selected as the maximum over several preprocessing schemes, DNN variants, and fit dimensionalities. The paper reports all variants, which is commendable, but it does not account for this post-hoc selection in the significance statement. At minimum, state that the 1.74 sigma is a post-hoc optimum and discuss the look-elsewhere effect, or pre-specify the PC3 fit as the analysis plan.","section":"Table I and Fig. 8"}],"minor_comments":[{"comment":"The title contains 'd eep learning' with a spurious space; it should read 'deep learning'.","section":"Title"},{"comment":"The phrase 'principle component analysis' is used repeatedly; the correct statistical term is 'principal component analysis'.","section":"Abstract and throughout"},{"comment":"The labels in Fig. 7 are garbled: the LL panel appears twice with '(a)' and the TL panel is also labeled '(b)' twice; the intended labels should be (a) LL, (b) TL, and (c) qqZZ.","section":"Fig. 7"},{"comment":"The legend entry 'DNN ( article-based)' should be 'DNN (particle-based)'.","section":"Fig. 1"},{"comment":"There are several typos: 'Prbing' should be 'Probing', 'gaugue boson' should be 'gauge boson', and 'Comparision' should be 'Comparison'.","section":"General text"},{"comment":"The PACS line lists keywords ('HL-LHC, vector boson scatter, electroweak') rather than PACS codes; either provide actual PACS codes or remove the line.","section":"PACS line"},{"comment":"Reference [29] is a Wikipedia article; a standard textbook or review reference for principal component analysis would be more appropriate.","section":"Reference [29]"},{"comment":"The paper does not state whether the code or trained models will be made available. A data-availability statement would improve reproducibility.","section":"Reproducibility"}],"recommendation":"major_revision","confidential_remarks":"This is a fast-simulation machine-learning feasibility study. The central methodological message, that a particle-based DNN with PCA on outputs can improve over a BDT for the LL fraction, is plausible and worth publishing after revision. The main obstacle is that the quantitative HL-LHC projection is built on no-pileup simulation and a flat systematic uncertainty; these are acknowledged but not validated. I would support acceptance once the pileup robustness is addressed or the claim is appropriately scoped, and once the fit procedure is specified in sufficient detail for the 1.74 sigma result to be reproducible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a sensible, clearly written feasibility study. It applies the same team's particle-based DNN from the WW analysis to the ZZ→4l channel, adds Yeo-Johnson preprocessing and PCA on the five output scores, and reports an expected significance of 1.74σ for the longitudinal fraction at 3000 fb⁻¹, up from the BDT baseline of 1.41σ. That is a modest but real gain, and the method of fitting multiple principal components from the DNN output is a legitimate extension worth knowing about.\n\nThe simulation pipeline is coherent: MadGraph+Pythia+Delphes, polarization samples handled separately, a clear event selection, and an Asimov significance calculation. The BDT baseline reproduces the published CMS projection of about 1.4σ, which gives me confidence that the comparison is honest. The paper does not oversell the physics; a 1.74σ expectation is still below discovery, and they correctly note that ATLAS+CMS could push it above 2σ.\n\nThe soft spots are real, though not fatal for the paper's narrow purpose. The stress-test note is right: pileup is neglected, and the DNN consumes exactly the objects pileup distorts—jets and isolated leptons. At 140–200 pileup interactions, the jet pair can change, the b-veto can fail differently, and lepton isolation degrades. Since the claimed improvement over the BDT is only 0.33σ, even a moderate pileup effect could erase most of it. The paper should at least quantify this, or the headline number should be labeled as a no-pileup upper bound.\n\nThe flat 10% systematic uncertainty on both signal and background is crude, but acceptable for a projection. More annoying is the selection of the best result among several PC combinations (PC1, PC1+PC2, PC1+PC2+PC3) without a trials correction; this introduces a small optimism bias. No public code or data is provided, which makes reproduction harder, but the setup is described well enough to be independently implemented.\n\nThis is not a groundbreaking paper, but it is an honest, useful incremental contribution for people working on VBS polarization or on ML-based multivariate fits in LHC analyses. The central argument holds within the no-pileup simulation, and the main caveat is checkable. Send it to peer review; the pileup question should be addressed in revision, and the multiple-comparison issue should be clarified.","headline":"A clean fast-simulation study showing a modest but real gain for longitudinal ZZ extraction via a particle-based DNN with PCA, yet the HL-LHC projection rests on a no-pileup assumption that could easily erase the improvement.","tokens_in":8814,"tokens_out":2123,"would_cite":true,"duration_ms":24223,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A particle-based deep neural network that preprocesses inputs with a power transformation and decorrelates its outputs with PCA raises the expected significance for measuring the longitudinal ZZ polarization fraction at the HL-LHC to 1.74…","keywords":["vector boson scattering","ZZ polarization","longitudinal fraction","deep neural network","principal component analysis","Yeo-Johnson transformation","HL-LHC","helicity fraction"],"falsifier":"Re-run the same Monte Carlo analysis with pileup included at 140 to 200 collisions per bunch crossing and compare the DNN-PCA significance with the BDT baseline; if the DNN-PCA significance drops to the BDT's 1.41 sigma or below, the claimed improvement is an artifact of the pileup-free simulation, while staying above about 1.6 sigma would confirm the central claim.","tokens_in":7755,"feed_emoji":"⚛️","tokens_out":10784,"duration_ms":92966,"temperature":0.7,"pith_summary":"This paper tries to establish that a deep-learning pipeline can measure the longitudinally polarized fraction of ZZ vector-boson scattering at the High-Luminosity LHC with an expected significance of about 1.7 standard deviations, using the full planned dataset of 3000 inverse femtobarns. The pipeline feeds the four-momenta of the four leptons and two jets into a per-particle neural network, preprocesses the inputs with a power transformation plus standardization, then applies principal-component analysis to the network's five class scores and fits the leading three components in three dimensions. This matters because the longitudinal (LL) component of vector-boson scattering is the piece tied directly to unitarity restoration through the Higgs mechanism and to possible new physics; a 1.7-sigma expectation would improve on the boosted-decision-tree baseline of about 1.4 sigma and would make a combined measurement by the two HL-LHC experiments likely to exceed 2 sigma.","feed_headline":"Longitudinal ZZ signal rises to 1.74 sigma with deep learning","feed_subtitle":"A particle-based network with PCA on its outputs beats the standard boosted-tree approach at the High-Luminosity LHC.","key_machinery":"The central object is a particle-based deep neural network: each lepton and jet enters as its own small input block built from its four-momentum, the blocks are merged hierarchically, and the network ends in five output nodes scoring the classes LL, TL, TT, qqZZ, and ggZZ. Two preprocessing steps carry much of the gain: the Yeo-Johnson power transformation, which makes skewed input distributions nearly Gaussian and handles negative values, followed by standardization computed on the LL training sample. The final mechanism is PCA on the five output scores; the leading three principal components, which together explain about 96% of the variance, are fitted simultaneously in three dimensions. This multidimensional fit extracts more signal than a cut on any single discriminant because the five class scores carry correlated shape information that PCA concentrates into a few decorrelated axes.","core_discovery":"The paper's central claim is that in the fully leptonic ZZ channel, the expected significance for the LL polarization fraction at 3000 inverse femtobarns can be raised from 1.41 sigma with a two-step boosted decision tree to 1.74 sigma with a particle-based DNN whose five class outputs are decorrelated by PCA and fitted in three dimensions; including a flat 10% systematic uncertainty on signal and background, the DNN-PCA still gives 1.66 sigma. The paper compares several discriminants and finds that the particle-based DNN with a power transformation plus standardization has the best single-classifier separation. It then shows that treating the five output scores as a vector, rotating them with PCA, and fitting the leading three principal components yields more significance than either a single discriminant or a two-step DNN. The paper presents this as evidence that this machine-learning pipeline makes the longitudinal VBS measurement feasible at the HL-LHC and transferable to other helicity-fraction measurements.","pith_inferences":["The PCA-on-outputs step suggests a general recipe for multiclass classifiers in high-energy physics: instead of thresholding one output score, keep the full score vector, decorrelate it, and fit the leading components; this can be tested on any multiclass tagger already in use.","If pileup at the HL-LHC degrades performance, the preprocessing is the first place to re-adapt: the power transformation is fit to the LL training sample, so a pileup-contaminated sample would shift both the transformation and the PCA axes.","A significance of 1.74 sigma still falls short of discovery; the method's real payoff may come from combining the ZZ channel with WW and WZ scattering, where the same longitudinal-fraction question has higher rates, or from using the PC distributions as inputs to a broader Higgs-coupling fit."],"forward_implications":["At 3000 inverse femtobarns the longitudinal ZZ fraction could be extracted with an expected significance near 1.7 sigma, better than the 1.41-sigma BDT baseline.","Combining the measurements of the two HL-LHC experiments would push the expected significance above 2 sigma, moving from sensitivity toward evidence for the longitudinal component.","Because the pipeline uses only per-particle four-momenta and multiclass scores, the same recipe can be applied to helicity-fraction measurements in other vector-boson-scattering channels.","Since the LL component grows strongly at high mass when Higgs or gauge couplings deviate from the standard model, the same DNN-PCA analysis could serve as a model-independent probe of Higgs couplings and anomalous gauge couplings."],"supporting_citations":[{"why":"introduces the particle-based DNN architecture that this paper adopts and enhances.","marker":"[14]"},{"why":"provides the BDT analysis and the 1.4-sigma expected significance that serve as the baseline to beat.","marker":"[15]"},{"why":"defines the Yeo-Johnson power transformation applied to every input variable before training.","marker":"[20]"},{"why":"supplies the fast detector simulation used to turn generated events into reconstructed leptons and jets.","marker":"[23]"},{"why":"sets the simulation configuration with pileup neglected that this analysis follows.","marker":"[24]"},{"why":"defines the likelihood-based significance calculation used to quote all expected significances.","marker":"[28]"},{"why":"demonstrates the earlier combination of a two-dimensional fit with a DNN in WW scattering that motivates the multidimensional fit here.","marker":"[13]"}],"fun_headline_variants":["Deep learning boosts ZZ polarization signal to 1.74 sigma","Particle DNN plus PCA lifts longitudinal ZZ to 1.74 sigma","Neural network achieves 1.74 sigma for longitudinal ZZ scattering","PCA on DNN outputs sharpens ZZ polarization to 1.74 sigma","Deep learning improves longitudinal ZZ significance to 1.74 sigma"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The projection relies on neglecting pileup—the detector simulation is configured for the HL-LHC but with the extra proton-proton collisions per bunch crossing turned off—so if pileup degrades lepton isolation or jet properties enough to reduce the network's separation power, the quoted 1.74-sigma significance would be optimistic.","fun_headline_variants_meta":{"raw":{"variants":["Deep learning boosts ZZ polarization signal to 1.74 sigma","Particle DNN plus PCA lifts longitudinal ZZ to 1.74 sigma","Neural network achieves 1.74 sigma for longitudinal ZZ scattering","PCA on DNN outputs sharpens ZZ polarization to 1.74 sigma","Deep learning improves longitudinal ZZ significance to 1.74 sigma"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000568,"raw_usage":{"total_tokens":2641,"prompt_tokens":849,"completion_tokens":1792,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":465,"completion_tokens_details":{"reasoning_tokens":1699}},"tokens_in":465,"tokens_out":1792,"duration_ms":10537,"temperature":1.0,"reasoning_tokens":1699,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:20:34.110738+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same Monte Carlo analysis with pileup included at 140 to 200 collisions per bunch crossing and compare the DNN-PCA significance with the BDT baseline; if the DNN-PCA significance drops to the BDT's 1.41 sigma or below, the claimed improvement is an artifact of the pileup-free simulation, while staying above about 1.6 sigma would confirm the central claim.","supporting_citations":[{"cited_title":"Dream Machines","cited_arxiv_id":"1808.06036","evidence_quote":"introduces the particle-based DNN architecture that this paper adopts and enhances."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines the Yeo-Johnson power transformation applied to every input variable before training."}],"review_version":1}