{"id":"e37ab42e-04e6-4799-99ba-d5b6d038aaa7","arxiv_id":"2412.00198","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Physics-inspired data augmentation halves the signal data requirement for CWoLa weak supervision searches, cutting the practical sensitivity threshold from roughly 6 sigma to roughly 3 sigma.","lead":"This paper demonstrates that data augmentation, specifically pT smearing and jet rotation, lowers the amount of signal data a weakly supervised neural network needs to beat a simple mass cut, from about 6 sigma to about 3 sigma in a simulated LHC search. The result offers a practical route to make model-agnostic anomaly searches more viable when signal samples are small.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Augmented pT smearing may reintroduce mjj dependence, and the Sec. 4.2 sculpting test does not cover it, so the reported threshold reduction could be inflated.","rationale":"The reader's weakest assumption correctly identifies the CWoLa independence requirement as load-bearing. My stress-test sharpens it: the augmentation pipeline changes the preprocessing, so the background-only flatness check in Sec. 4.2 is no longer sufficient. The pT smearing step has a pT-dependent noise width, and because SR and SB differ in typical pT/mjj, the smeared and re-normalized images can acquire an mjj dependence that the classifier could exploit. Since signal and background in the SR also have different mjj distributions, such an artifact would directly inflate the sensitivity gain that the paper attributes to learning signal substructure. This is not merely a stylistic or reproducibility concern; it challenges the internal validity of the headline number. I did not find a more fundamental internal inconsistency: the network architecture, sample preparation, and significance formulas are standard, and the qualitative improvement from augmentation is plausible. The typo in Eq. (5.2) and the use of 1% rather than 5% systematic uncertainty are secondary. The proposed augmented sculpting check would settle whether the concern lands, so the reader's conditional verdict stands, with this check added as a prerequisite.","tokens_in":12997,"tokens_out":6340,"duration_ms":64456,"concrete_test":"Repeat the Sec. 4.2 background-only sculpting test using the +5 pT-rot augmentation and EN normalization: train only on background SR/SB samples, then plot the NN-cut efficiency ε(mjj) at εSB = 10%, 1%, and 0.1%. If ε(mjj) is not flat within the available statistics, the augmentation reintroduces sculpting and the reduced learning threshold in Figs. 5-7 is not a pure signal-substructure effect. Use the same number of augmented copies as in the headline results and enough MC events to resolve a 10% slope.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that data augmentation lowers the CWoLa learning threshold from ~6σ to ~3σ. This is only meaningful if the trained network is learning signal jet substructure rather than the SR/SB boundary. CWoLa requires the normalized jet-image inputs to be mjj-independent after all preprocessing. The paper's background-only sculpting test (Sec. 4.2) validates this only for the un-augmented pipeline. The augmentation pipeline used for the headline results, especially pT smearing (Eq. 5.1), is applied before preprocessing/EN normalization. The Gaussian smearing width f(pT) depends on pT, so the noise added to high-pT constituents differs from low-pT constituents. SR and SB events have different typical mjj and hence different pT scales, so this pT-dependent noise can survive per-event EN standardization and create a residual mjj dependence in the augmented image distribution, even when the un-augmented images are mjj-independent. Jet rotation alone does not have this issue, but the combined pT-rot method, which gives the best reported performance, does. If the augmented classifier learns this mjj dependence, it can separate the SR and SB training sets, and because signal and background within the SR have different mjj distributions, the ROC-based sensitivity in Figs. 5-7 would be inflated. No augmented sculpting check is reported, so the headline reduction rests on an unvalidated CWoLa assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes physics-inspired data augmentation as a way to reduce the amount of signal data needed for weakly supervised (CWoLa) searches. Using a simulated Hidden Valley dijet resonance benchmark at 13 TeV, the authors compare pT smearing, jet rotation, and their combination against the un-augmented pipeline in terms of ROC-based sensitivity curves with error bars from ten retrainings. They report that the learning threshold is reduced from roughly 6σ to 3σ, that the combined method performs best, that performance saturates as the number of augmented copies grows, and that the gains survive a 1% background systematic uncertainty.","tokens_in":13389,"tokens_out":5779,"duration_ms":59470,"significance":"If valid, the result would be a simple and practical improvement for resonant anomaly searches: it lowers the data requirement of CWoLa without changing the search strategy or relying on external simulation labels. The paper has several strengths: it uses a concrete benchmark, reports mean and standard deviations from repeated training, compares with full supervision, and explicitly discusses limiting behavior. The central technical risk is that the augmentation may break the mjj-independence assumption on which CWoLa rests, and the manuscript does not currently rule this out; the headline threshold reduction is therefore not yet fully established.","major_comments":[{"comment":"The sculpting check in Sec. 4.2 is performed only on the un-augmented jet-image pipeline. The pT smearing of Eq. (5.1) is applied to constituents before preprocessing and event normalization, and the smearing width f(pT) depends on pT. Since SR and SB events have different mjj distributions and hence different jet pT spectra, the augmented feature distributions can in principle retain a residual mjj dependence even when the un-augmented images are mjj-independent. If the classifier exploits that dependence, it can separate the SR and SB training sets and inflate the ROC-based sensitivities in Figs. 5–7. Because the best-performing method (pT-rot) includes pT smearing, the central claim rests on this untested assumption. Please repeat the background-only sculpting test of Sec. 4.2 for the augmented pipelines (at least +5 and +20 pT-rot), and report whether the mjj distributions and selection efficiencies are flat; if sculpting appears, quantify how much of the threshold reduction survives after removing it.","section":"Sec. 4.2 and Sec. 5.1 (Eq. 5.1)"},{"comment":"The headline claim of a reduction from roughly 6σ to 3σ relies on the notion of a learning threshold, but the paper gives only a qualitative definition in the Introduction and appears to read thresholds from the crossing of curves in the figures. Please specify the exact extraction rule used to define the threshold, including whether the crossing is evaluated on the mean curve or on individual retrainings, how interpolation is performed, and how the standard deviation is propagated. Without such a rule, the central quantitative claim is not precisely reproducible.","section":"Sec. 4.3 and Figs. 3, 5–8"}],"minor_comments":[{"comment":"Equation (5.2) appears to contain a typo: the second coordinate should presumably read η′ sin θ + φ′ cos θ rather than η′ sin θ + φ′ sin θ.","section":"Sec. 5.1, Eq. (5.2)"},{"comment":"The systematic-uncertainty study uses a relative background uncertainty of 1%, while the text notes that the typical value is 5%; since the compression is stronger at larger uncertainty, the statement that augmentation remains beneficial under systematic uncertainty would be better supported by also showing the 5% case.","section":"Sec. 5.4, Fig. 8"},{"comment":"The statistical treatment of the repeated trainings could be clarified: it is not stated whether the 10 retrainings reuse the same test sample, and how much of the quoted standard deviation is due to training variability versus finite test statistics.","section":"Sec. 3.2"},{"comment":"For pT smearing, the manuscript does not state how negative resampled pT values are handled; a brief sentence on clipping or truncation would remove ambiguity.","section":"Sec. 5.1"},{"comment":"No code or data release is mentioned; providing the training and preprocessing pipeline would strengthen reproducibility of the numerical claims.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of the journal and the central idea is interesting, but the missing augmented sculpting test is a genuine validity risk for the main quantitative claim. If the authors add that test and it is clean, I would expect the paper to be publishable; if the augmented pipeline shows mjj sculpting, the conclusions will need to be substantially revised. The citation pattern is appropriate and I see no disclosure concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper shows that physics-inspired data augmentation (jet rotation and pT smearing) can cut the training-data threshold for CWoLa weak supervision roughly in half, from ~6σ to ~3σ, on a simulated Hidden Valley benchmark. The result is plausible and the experiments are clean, but there's one load-bearing gap: the sculpting test that validates the CWoLa assumption is only done for the un-augmented pipeline. Since pT smearing is applied before the mjj-normalization, it could in principle reintroduce mjj dependence, and the best performing method includes that smearing. The paper doesn't check this, so the headline improvement might be inflated by the classifier learning the SR/SB boundary instead of jet substructure.\n\nWhat's new: I believe this is the first quantitative study of data augmentation specifically for CWoLa. The asymptotic analysis (diminishing returns with more augmented copies) is useful. The paper is also careful to show error bars from 10 retrainings and to report that η-ϕ smearing and Gaussian noise didn't help.\n\nWhere the soft spots are, in order of severity:\n\n1. The missing augmented sculpting test. Section 4.2 validates EN normalization on background-only samples without augmentation. The best pipeline (pT smearing + rotation) applies pT smearing before preprocessing, and the smearing width has a pT-dependent form (Eq 5.1). Since SR and SB have different typical mjj and therefore different jet pT scales, the smeared images might differ between the two regions in a way that survives event normalization. A background-only test with exactly the augmented pipeline would settle this. If the augmented classifier can separate SR from SB on background, the sensitivity estimates in Figs 5-7 are not trustworthy. This is my main issue.\n\n2. The systematic uncertainty curves use 1% relative background uncertainty, while the text notes 5% is typical. Showing 1% is a bit optimistic; the qualitative conclusion probably survives with 5%, but it would be better to show it.\n\n3. Minor: Eq. (5.2) has a typo in ϕ″ (second term should be ϕ′ cos θ). No code/data is released, which makes the exact augmentations hard to reproduce.\n\nThe paper is worth a serious referee. The central idea is sound and the experiment is well-designed; it just needs the augmented sculpting check and a few smaller fixes. I'd probably cite it once the check is done, but not yet. This paper is for the anomaly detection / weak supervision crowd and for LHC analysts who need to know how much signal is needed to make CWoLa work.\n\nRecommendation: send to peer review, but the authors must add a background-only sculpting test using the augmented pipeline before this is accepted.","headline":"Useful empirical result, but the missing augmented sculpting check could undermine the headline threshold reduction.","tokens_in":13786,"tokens_out":4627,"would_cite":false,"duration_ms":42577,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Data augmentation halves the signal data needed for weak-supervision searches.","keywords":["weak supervision","data augmentation","CWoLa","jet images","physics-inspired augmentation","Hidden Valley model","learning threshold","dijet resonance search"],"falsifier":"Train the augmented CWoLa classifier on background-only data across the same signal and sideband regions, applying the same pT-smearing and jet-rotation augmentation. If the NN cut efficiency shows a statistically significant dependence on mjj beyond Monte-Carlo fluctuations, the augmentation breaks the mjj-independence assumption and the reported sensitivity gains may be partly spurious.","tokens_in":12784,"feed_emoji":"🧠","tokens_out":9647,"duration_ms":73746,"temperature":0.7,"pith_summary":"Weak supervision trains a classifier on mixed signal-enriched and signal-depleted samples, but it typically needs so many signal events that it is not competitive with traditional searches. This paper asks whether physics-inspired data augmentation—smearing jet transverse momenta to mimic detector resolution and rotating jet images by random angles—can lower that requirement. Using a Hidden Valley benchmark with jet-image inputs, the authors report that adding five augmented copies per event reduces the learning threshold from about 6σ to about 3σ in both benchmark scenarios, and cuts the training-to-training variance in half. The combined augmentation method works best, and the improvement survives a 1% background systematic uncertainty.","feed_headline":"Data augmentation halves weak-supervision learning threshold","feed_subtitle":"pT smearing and jet rotation cut needed signal from about 6σ to about 3σ.","key_machinery":"The working mechanism is label-preserving data augmentation applied to jet images before pixelation. pT smearing resamples each jet constituent's transverse momentum from a normal distribution whose width is the detector energy-resolution function, $\\sqrt{0.05^2 p_T^2 + 1.50^2 p_T}$ (with $p_T$ in GeV), so the network sees realistic detector fluctuations; jet rotation rotates the (η, φ) coordinates of each jet about its center by an independent random angle in [−π, π], exposing the network to the rotational diversity of jet substructure. These transformations enlarge the training set and force the classifier to learn signal/background differences that are robust to detector effects and orientation. The paper also relies on event normalization of the jet images to remove $m_{jj}$ dependence from the inputs, which prevents the classifier from sculpting a fake signal by learning the region definitions.","core_discovery":"The central claim is that data augmentation converts a weak-supervision classifier from a data-hungry method into a practical one. In the CWoLa framework, a neural network is trained on dijet events from a signal region and a sideband defined by the dijet invariant mass, each containing a mixture of signal and QCD background. The paper shows that by augmenting the training jet images with pT smearing (resampling constituent momenta with the detector resolution function) and jet rotation (rotating each jet by an independent random angle in [−π, π]), the network reaches the same post-cut sensitivity with roughly half the pre-cut signal sensitivity. Specifically, the learning threshold drops from around 6σ to around 3σ for both the indirect-decay and direct-decay Hidden Valley scenarios, and the standard deviation of the sensitivity across 10 retrainings is reduced to about half. The combination of both augmentations outperforms either alone, and the benefit persists when a 1% relative background systematic uncertainty is included.","pith_inferences":["The same augmentation recipe could likely be adapted to other weak-supervision or anomaly-detection setups that use jet images, such as autoencoder-based outlier detection, as long as the transformation preserves the label.","The rotation augmentation effectively builds rotational invariance into the network without adding a custom architecture, implying that other known symmetries of jet images might be injected similarly.","Since the paper's benchmark is a single resonance topology, a natural testable extension is whether the threshold reduction holds for multi-pronged or boosted-object signals, where rotation diversity may matter even more.","Combining data augmentation with pre-training or transfer-learning approaches could push the learning threshold even lower; this stacking is not tested here."],"forward_implications":["Weak-supervision searches can be carried out with roughly half the integrated luminosity or signal yield previously needed, making them viable earlier in an LHC run.","Physics-informed augmentation that respects detector resolution and rotational symmetry is a general lever for improving classifier generalization in collider anomaly detection.","The reduction in training variance means a single trained network is more reliable, so searches can trust the sensitivity estimate from one training run.","Because performance saturates after about +30 augmented copies, practitioners can choose a modest augmentation factor and avoid diminishing returns from larger datasets.","The method's robustness to a 1% background systematic suggests it can be combined with standard background-estimation uncertainties in a real search."],"supporting_citations":[{"why":"Introduces the classification-without-labels method that trains a classifier on mixed signal and background samples.","marker":"[5]"},{"why":"Establishes the hunting setup that defines signal and sideband regions from a resonant variable to build mixed training datasets.","marker":"[7]"},{"why":"Provides the benchmark scenarios and the learning-threshold concept that this paper extends to data augmentation.","marker":"[11]"},{"why":"Supplies the underlying physics model that produces the signal jets with dark showers.","marker":"[19]"},{"why":"The detector simulation used to define the energy resolution for the pT-smearing augmentation.","marker":"[26]"},{"why":"Inspires the physics-based augmentation strategy by demonstrating that label-preserving physical transformations improve learning.","marker":"[36]"}],"fun_headline_variants":["Augmentation halves signal needed for weak supervision","pT smearing and jet rotation cut weak-supervision threshold","Weak-supervision signal requirement drops from 6σ to 3σ","Data augmentation halves weak-supervision data hunger"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method relies on the premise that, after event normalization, the jet-image features are statistically independent of the dijet invariant mass mjj; if residual mjj dependence survives or is introduced by augmentation, the classifier could fake a signal by learning the region definitions.","fun_headline_variants_meta":{"raw":{"variants":["Augmentation halves signal needed for weak supervision","pT smearing and jet rotation cut weak-supervision threshold","Weak-supervision signal requirement drops from 6σ to 3σ","Data augmentation halves weak-supervision data hunger"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000174,"raw_usage":{"total_tokens":1226,"prompt_tokens":830,"completion_tokens":396,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":446,"completion_tokens_details":{"reasoning_tokens":329}},"tokens_in":446,"tokens_out":396,"duration_ms":4178,"temperature":1.0,"reasoning_tokens":329,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:37:42.863864+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the augmented CWoLa classifier on background-only data across the same signal and sideband regions, applying the same pT-smearing and jet-rotation augmentation. If the NN cut efficiency shows a statistically significant dependence on mjj beyond Monte-Carlo fluctuations, the augmentation breaks the mjj-independence assumption and the reported sensitivity gains may be partly spurious.","supporting_citations":[],"review_version":1}