{"id":"b40230e9-bd42-4a2d-8101-7a67f763d468","arxiv_id":"2501.01778","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"NPLM-based classifiers and end-to-end NPLM outperform BDT-based anomaly detection at low signal injection on the LHCO and RODEM benchmarks, with lower variance across hyperparameters.","lead":"This paper tests whether a machine-learning method called NPLM can beat standard boosted decision trees at spotting rare new-physics resonances in collider data. It reports that NPLM-based pipelines are more sensitive at very low signal rates and give more stable results across hyperparameter choices, though only under an idealized perfect-background assumption.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim rests on a perfect background-template assumption that is untested for the NPLM-classifier's advertised use case; a misspecified template could remove or reverse the reported advantage.","rationale":"The reader's weakest_assumption correctly identifies the perfect-template limitation and the lack of background uncertainty simulation. The paper explicitly acknowledges the idealized setting, but the NPLM-classifier approach is motivated by the opposite scenario, making this the most load-bearing gap in the argument. If an imperfect template breaks the calibration or erases the sensitivity gain, both components of the central claim (better detection and lower variance) lose support in the realistic regime the paper aims to address. The reader's CONDITIONAL verdict already reflects this concern, so no change is needed. A secondary concern about the asymmetric hyperparameter treatment of BDT versus NPLM further weakens the variance claim, but it is less fundamental than the template issue. The proposed test directly probes the advertised use case and would settle whether the central claim extends beyond the idealized benchmark.","tokens_in":12545,"tokens_out":6966,"duration_ms":73179,"concrete_test":"Re-run the NPLM-classifier vs BDT-classifier comparison of Sec. 4.3 with a misspecified background template in the signal region: for instance, construct R by fitting a smooth polynomial to the sidebands of mJJ and extrapolating into the signal region, then use that R for both training and calibration. Calibrate p-values with pseudo-experiments drawn from the same misspecified R. Compare the false-positive rate at a nominal alpha (e.g., 0.05) and the discovery power at N(S)/N(B) = 8.2e-4 against the identical BDT pipeline. If the NPLM false-positive rate exceeds the nominal value or the power advantage reverses, the central claim fails outside the idealized setting.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims NPLM-based methods outperform BDT-based approaches, especially at low signal injection. The numerical evidence for the NPLM-classifier (Sec. 4.3, Figs. 3-5, Table 1) is obtained exclusively in the idealized setting where the background template R is perfect, as stated in Sec. 4.2: 'R in this work pertains to the idealised setting.' However, the NPLM-classifier is introduced in Sec. 3.1 precisely for the case 'in absence of a good background modelling.' The paper never tests that case. If R is biased relative to the true background in the signal region, NPLM's in-sample training will fit the bias as if it were signal. The cut-and-count statistic (Eq. 7) then becomes inflated even under the null hypothesis, and the calibration in Eq. 6 uses pseudo-experiments generated from the same biased R, so the resulting p-values are not valid. The reported power advantage at low signal injection could be an artifact of a perfect template; under realistic template misspecification the advantage may vanish or reverse. Since approach (2) is motivated by imperfect background modelling, the central claim is not supported for its intended domain.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares NPLM-based resonant anomaly detection strategies with the standard BDT-based CWoLa approach on the LHCO and RODEM benchmark datasets. Two NPLM use cases are studied: (1) an end-to-end NPLM test over all six input features, applied when an accurate background template is available, and (2) an NPLM-based classifier used to select events in the signal region, followed by a cut-and-count test, including a hyper-test over several selection thresholds. Detection power is reported as power curves and median Z-scores at three signal-injection levels, with the central claims that NPLM outperforms BDTs at low signal injection and that NPLM reduces hyperparameter-induced epistemic variance.","tokens_in":12787,"tokens_out":5797,"duration_ms":61129,"significance":"The paper addresses a timely and practically relevant question: how to make resonant anomaly detection more sensitive to rare signals and more stable under hyperparameter choices. Its strengths are the use of public, independent benchmark datasets (LHCO and RODEM), empirical calibration via signal-free pseudo-experiments, a transparent hyperparameter scan on the BDT side, and a multiple-testing procedure over selection thresholds that reduces variance. The numerical evidence is internally consistent in the idealized perfect-template setting. However, the main comparison for the NPLM-classifier, which is motivated for the case of absent or imperfect background modelling, is carried out only under a perfect background template; the paper does not test how either method behaves under template misspecification. The variance comparison is also not fully symmetric in the hyperparameter spaces considered. If the missing misspecification study can be added and the variance claim appropriately qualified, the paper will be a solid contribution to the anomaly-detection literature.","major_comments":[{"comment":"The NPLM-classifier is introduced in Section 3.1 for use 'in absence of a good background modelling,' and Section 5 repeats that the selection-plus-calibration strategy is the appropriate one when the template is not accurately known. Yet all numerical experiments for this approach assume the idealized setting in which the template R exactly reproduces the background distribution in the signal region, as stated in Section 4.2: 'R in this work pertains to the idealised setting.' This is a load-bearing gap: with a biased template, the in-sample NPLM fit will absorb the template bias into f_w(x), the cut-and-count statistic in Eq. (7) will be shifted even under the null hypothesis, and calibration pseudo-experiments generated from the same biased R will not yield valid p-values. The reported low-signal advantage in Figs. 3 and 5 and Table 1 may therefore not transfer to the intended application. Please repeat the central comparison under at least one realistic misspecification, for example a smooth sideband fit with a shape bias in m_JJ or a template built from a shifted signal-region definition, and report null-calibration and power for both NPLM-classifier and BDT-classifier under the same misspecified R. If such a study cannot be included, the claims in Section 5 for approach (2) should be substantially weakened.","section":"§3.1, §4.2, §4.3"},{"comment":"The claimed reduction in epistemic variance is not an apples-to-apples comparison. For BDTs, the hyper-test is applied only over the selection threshold thr, while the NPLM hyper-test inherits the multiple-testing procedure over the kernel width sigma from Ref. [23] and reports the remaining spread over the other NPLM hyperparameters. Table 1 and Fig. 5 therefore compare a BDT variance that includes sensitivity to nleaf and lambda but not to the hyper-test over those parameters, with an NPLM variance that has already been partially reduced by multiple testing over sigma. The paper itself acknowledges in Section 4.4 that extending the BDT hyper-test to multiple BDT hyperparameters is left to future work. Please either implement that extension, or restrict the claim to a statement about the specific hyperparameter sets used here and state explicitly how many models contribute to each shaded band.","section":"§4.4, Table 1"},{"comment":"The central low-injection advantage rests on small absolute differences in median Z-score: at N(S)/N(R) = 8.2 x 10^-4, Table 1 reports 0.30 +/- 0.04 for NPLM 5D + hyper-tcc versus 0.17 +/- 0.04 for BDT 5D + hyper-tcc. The paper does not state the number of pseudo-experiments used in Figs. 1-5 or the correlation structure between the two procedures, so it is unclear whether the difference is statistically significant rather than a fluctuation in the toy ensemble. Please report the number of toys, and provide paired bootstrap confidence intervals or an equivalent measure for the power curves and for the median Z-scores in Table 1.","section":"§4.3, Table 1"}],"minor_comments":[{"comment":"The definition of Z_alpha in Eq. (8) is phrased as the quantile of the normal distribution 'at the alpha complement to 1'; please write it explicitly as z_alpha = Phi^{-1}(1 - alpha) and define what is plotted on the horizontal axis of the power curves.","section":"§4.2, Eq. (8)"},{"comment":"The table caption says 'average Z-score' while the text repeatedly says 'median Z-score' and the main text says 'median Z-score among different values of the tunable hyperparameters.' Please make this consistent, and clarify whether the reported uncertainty is the standard deviation of the median or the standard error of the mean.","section":"Table 1 and captions"},{"comment":"The captions use 'standard deviation' in some places and 'standard error' in others for the shaded bands; please choose one consistent definition and state it in a common caption note.","section":"Figs. 3-5"},{"comment":"Equation (7) subtracts the number of selected template events from the number of selected data events and divides by the square root of the selected template count. Please state explicitly whether R is a fixed reference sample or is resampled in the pseudo-experiments, and whether the finite size of R is accounted for in the denominator.","section":"§4.2, Eq. (7)"},{"comment":"There are several typographical and grammatical issues, including 'a end-to-end,' 'subjettinness,' 'F ALKON' with an internal space, and 'hyperparameters choice.' A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The paper mentions the look-elsewhere effect when scanning the resonant variable but does not compute a global p-value for the sliding-window scan. Please clarify whether the reported results are local p-values only and whether global p-value combination is intended as future work.","section":"§3.1 and §5"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent about its idealized-setting assumption, which is commendable, but that assumption is not merely a technical convenience: it is the regime where the NPLM-classifier is claimed to be beneficial for the case of imperfect background modelling. The missing misspecification study is the main reason for my recommendation. The self-citation pattern ([16,17,23]) is understandable since the paper builds directly on that prior work, and the central comparison uses independent public benchmarks. The scope is appropriate for a specialized hep-ex or ML-for-science venue; the paper does not yet meet the bar for a broader journal because the headline claim about 'robust' anomaly detection is not supported beyond the perfect-template setup."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper does something real—it hooks NPLM into the standard sliding-window resonant search and shows, on LHCO and RODEM, that it beats the BDT+CWoLa baseline at low signal injection and cuts hyperparameter variance. The numbers are plausible and the authors are transparent that all of it is in the idealized perfect-template setting. The central claim for that setting holds up. Two soft spots keep me from signing on to the broader robustness story.\n\nFirst, the paper motivates approach (2) (NPLM as classifier) for the case where accurate background modelling is unavailable, but every numerical test assumes a perfect template R in the signal region (Sec. 4.2). If R is biased, the in-sample training will fold that bias into the score, and the cut-and-count statistic in Eq. 7 is calibrated on pseudo-experiments drawn from the same biased R, so the p-values aren't valid. The authors acknowledge this implicitly by calling it idealized, but they never test how the advantage degrades under template misspecification. That is the gap between what the abstract claims (\"robust\") and what is actually demonstrated.\n\nSecond, the variance comparison is asymmetric. NPLM gets the multiple-testing treatment over kernel width sigma (following their prior [23]), while BDT hyperparameters are not included in the hyper-test; only thr is. So part of the \"reduced epistemic variance\" is built into the protocol. The authors note this and leave it to future work, which is fair, but Table 1's headline numbers should be read with that caveat.\n\nMinor: no code or data release, and the hyperparameter grids are small, but that's not disqualifying.\n\nFor a reader who wants to see NPLM in a bump-hunt pipeline, this is a useful, clearly written benchmark. For a reader who wants evidence that NPLM-classifiers survive realistic background uncertainty, the paper doesn't yet provide it. I'd send it to peer review—the idealized benchmark is worth publishing, and the missing pieces are addressable rather than fatal.","headline":"Solid integration of NPLM into resonant bump hunts, with honest caveats; the main advertised regime (imperfect background) is untested.","tokens_in":13336,"tokens_out":1892,"would_cite":true,"duration_ms":18550,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that in-sample density-ratio learning with NPLM detects rare resonant signals more reliably than BDT-based classifiers, and that a hyper-test over selection thresholds removes most of the variance from hyperparameter…","keywords":["resonant anomaly detection","NPLM","New Physics Learning Machine","boosted decision trees","CWoLa","bump hunt","likelihood-ratio test","multiple testing"],"falsifier":"Run the same LHCO benchmark with a background template produced by fitting the sidebands (rather than using the true background) and compare median Z-scores at the lowest signal fraction, $N(S)/N(B)=8.2\\times10^{-4}$; the paper's central claim fails if the NPLM median Z-score falls below the BDT hyper-test median.","tokens_in":12321,"feed_emoji":"🔬","tokens_out":7708,"duration_ms":67166,"temperature":0.7,"pith_summary":"The paper claims that the New Physics Learning Machine (NPLM) improves resonant anomaly detection in particle physics compared to the standard Boosted Decision Tree (BDT) approach, especially when the injected signal is rare. NPLM trains a kernel-based classifier on the full dataset and evaluates it in-sample, estimating the log-density ratio between data and a reference background template, rather than splitting the data into training and validation folds. This in-sample strategy keeps sensitivity to tiny signal admixtures, which the authors argue is why NPLM beats BDTs at the lowest signal fraction they test. The paper also shows that running NPLM end-to-end on all variables removes the variance induced by choosing a selection threshold, and that when only an imperfect background template is available, using NPLM as a classifier with a hyper-test across threshold values restores stability.","feed_headline":"In-sample NPLM beats BDTs at rare-signal detection","feed_subtitle":"A hyper-test over thresholds makes the hunt for rare resonances both more sensitive and more stable under hyperparameter choices.","key_machinery":"The central object is the NPLM training loss $L_{\\mathrm{NPLM}}[f_w] = \\sum_{x \\in R} w_x(e^{f_w}-1) - \\sum_{x \\in D} f_w$, where $R$ is the reference background sample and $D$ is the data sample; minimizing it produces a function $f_w$ that approximates the log-density ratio $\\log n(x|D)/n(x|R)$, and the test statistic is $t_{\\mathrm{NPLM}} = -2\\min_w L_{\\mathrm{NPLM}}[f_w]$. Because the classifier is trained and evaluated on the same full dataset, no events are held out, which preserves sensitivity to rare signal events. The end-to-end version feeds all six variables, including the resonant mass, into this loss and directly produces a Neyman-Pearson test statistic with no selection threshold. The classifier version instead computes a cut-and-count statistic after a threshold, and the paper stabilizes it with a hyper-test that takes the minimum p-value over several threshold choices, borrowing the multiple-testing idea from the recent NPLM literature.","core_discovery":"The central claim, stated on the paper's own terms, is that NPLM-based strategies outperform BDT-based classifiers in detection power at low signal injection while substantially reducing epistemic variance due to hyperparameter choices. The paper supports this with controlled numerical experiments on two benchmarks: the LHCO dijet dataset, where the signal is a narrow resonance, and the RODEM dataset, where the signal is a flat excess in the invariant mass. In the low-injection regime, the median Z-score of the NPLM-classifier with a hyper-test is up to a factor of three higher than the BDT counterpart, and the spread across hyperparameter choices is smaller. The authors attribute the gain to the in-sample nature of the NPLM test, which uses the full dataset, in contrast to the k-fold out-of-sample evaluation used for BDTs. The discovery is therefore a concrete demonstration that in-sample likelihood-ratio estimation is a useful alternative to standard classification for rare-signal searches.","pith_inferences":["Because the paper's numerical comparisons use a perfect background template, a natural next test is to replace it with a sideband-derived template and see whether NPLM's low-injection advantage survives real background uncertainty.","A testable prediction of the paper's explanation is that template errors hurt NPLM more than BDT pipelines, since the NPLM loss assumes the reference is the true no-signal density.","The hyper-test idea could be extended to jointly vary BDT hyperparameters as well as thresholds, which the paper itself leaves for future work.","Applying NPLM to a full LHC-style analysis with realistic statistical and systematic uncertainties would show whether the factor-of-three median Z-score improvement translates outside the idealized setting."],"forward_implications":["Searches for rare resonances should gain discovery power if NPLM replaces BDT classifiers in the anomaly-selection stage.","End-to-end NPLM removes the selection-threshold hyperparameter entirely, making the analysis less dependent on analyst choices.","The hyper-test over thresholds gives a stable way to combine classifier scores when the background template is imperfect.","The in-sample principle suggests that other full-data density-ratio learners could beat k-fold classifiers in low-signal regimes."],"supporting_citations":[{"why":"Introduces the NPLM algorithm that estimates the log-density ratio and provides a Neyman-Pearson test statistic.","marker":"[15]"},{"why":"Provides the neural-network NPLM implementation and the heuristic used to select well-behaved models.","marker":"[16]"},{"why":"Supplies the kernel-method NPLM implementation used in this paper's numerical experiments.","marker":"[17]"},{"why":"Contributes the multiple-testing procedure that removes threshold and hyperparameter selection variance.","marker":"[23]"},{"why":"Establishes BDTs as the strong weakly-supervised baseline that the paper compares against.","marker":"[14]"},{"why":"Introduces CWoLa, the weakly-supervised classification framework underlying the BDT approach.","marker":"[12]"},{"why":"Provides the LHCO dijet dataset used for the resonant benchmark and signal injections.","marker":"[24]"},{"why":"Provides the RODEM jet dataset used for the non-resonant benchmark.","marker":"[29]"}],"fun_headline_variants":["NPLM triples sensitivity in rare-resonance searches","Rare-signal detection: NPLM beats BDTs with less tuning","In-sample NPLM outperforms BDTs at low signal injection","NPLM cuts epistemic variance in anomaly detection","Robust NPLM finds rare resonances better than BDTs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The numerical comparisons assume a perfect background template in the signal region, so the reference sample exactly represents the no-signal distribution; if real templates carry substantial errors, the claimed advantage of in-sample training is not guaranteed.","fun_headline_variants_meta":{"raw":{"variants":["NPLM triples sensitivity in rare-resonance searches","Rare-signal detection: NPLM beats BDTs with less tuning","In-sample NPLM outperforms BDTs at low signal injection","NPLM cuts epistemic variance in anomaly detection","Robust NPLM finds rare resonances better than BDTs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000198,"raw_usage":{"total_tokens":1359,"prompt_tokens":930,"completion_tokens":429,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":546,"completion_tokens_details":{"reasoning_tokens":338}},"tokens_in":546,"tokens_out":429,"duration_ms":4463,"temperature":1.0,"reasoning_tokens":338,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:20:35.472768+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same LHCO benchmark with a background template produced by fitting the sidebands (rather than using the true background) and compare median Z-scores at the lowest signal fraction, $N(S)/N(B)=8.2\\times10^{-4}$; the paper's central claim fails if the NPLM median Z-score falls below the BDT hyper-test median.","supporting_citations":[{"cited_title":"Multiple testing for signal-agnostic searches of new physics with machine learning","cited_arxiv_id":"2408.12296","evidence_quote":"Contributes the multiple-testing procedure that removes threshold and hyperparameter selection variance."},{"cited_title":"Official Datasets for LHC Olympics 2020 Anomaly Detection Challenge (Version v6), 2019","cited_arxiv_id":null,"evidence_quote":"Provides the LHCO dijet dataset used for the resonant benchmark and signal injections."},{"cited_title":"RODEM Jet Datasets","cited_arxiv_id":"2408.11616","evidence_quote":"Provides the RODEM jet dataset used for the non-resonant benchmark."}],"review_version":1}