{"id":"79551739-2e43-4e90-9b87-d3f636591840","arxiv_id":"1908.08923","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Anomalous Higgs couplings change the shape of the di-Higgs mass distribution, and an unsupervised clustering algorithm captures those shape differences more finely than a hand-defined taxonomy.","lead":"This paper classifies the shapes of Higgs pair invariant mass distributions, computed at next-to-leading order with full top quark mass dependence, and maps each shape class to regions of the anomalous-coupling parameter space. It then applies an unsupervised machine-learning method (autoencoder plus KMeans clustering) to identify shape classes and proposes seven new benchmark points for future experimental analyses.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The maps and benchmark points inherit a 2D-slice sampling assumption: with other couplings fixed to SM values, simultaneous multi-coupling deviations are never explored.","rationale":"I read the paper as a proof-of-concept shape taxonomy built on an established NLO parametrization (Ref. [71]), and that underlying calculation is independent support for the numerical input. The load-bearing weakness is not the machine-learning machinery itself but the sampling used to derive the global coupling-shape maps and the benchmark points. The 2D projection is explicitly described, but the inference from those slices to five-dimensional coupling-shape associations is never tested. Eq. (2.5) contains cross terms involving three or more couplings, so the 2D slices could miss phenomenologically relevant regions of the 5D parameter space. The reader's weakest assumption points to exactly this issue, and I agree with the resulting CONDITIONAL verdict: the concern is concrete and addressable, not a fundamental error. A full 5D resampling would settle whether the maps and benchmarks survive, so I would keep the verdict unchanged rather than moving it.","tokens_in":17975,"tokens_out":7365,"duration_ms":77370,"concrete_test":"Use the A_i tables from Ref. [71] and Eq. (2.5) to generate a full 5D sample (e.g., 10^5-10^6 uniform or Latin-hypercube points over the ranges in Eq. (2.6)), apply the same autoencoder plus KMeans pipeline, and compare the resulting cluster centers and parameter maps with Figs. 10-16. Then apply the Section 3.3 benchmark-selection procedure to this full 5D sample and compare with Table 2. If new shape clusters appear, or if the ctt-shape association changes for a non-negligible fraction (say >10%) of the full-5D points satisfying the 6.9 sigma_SM cross-section limit, the 2D-slice results are not representative.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.3 states that in every figure where two couplings are varied, the other three are set to their SM values, and Section 3.2's cluster maps are drawn on the same ten 2D planes. The autoencoder input (Section 3.1) is described as a set of 10^5 distributions, but no full 5D sampling is specified, and the benchmark selection in Section 3.3 searches the 'input grids'—i.e., the 2D-slice grids. The central claim that the unsupervised method captures shape features and maps them onto the five-dimensional coupling space therefore depends on the implicit assumption that these 2D slices are representative of the full 5D space. This is not established. The differential cross section in Eq. (2.5) contains mixed terms involving three or more couplings (e.g., A9 c_tt c_ggh c_hhh, A17 c_t c_tt c_ggh, A18 c_t c_ggh^2 c_hhh), so simultaneous deviations can produce shapes that never appear in any slice where the remaining couplings are SM-valued. In particular, the conclusion that small nonzero ctt values are likely to produce a doubly peaked structure is based on slices where chhh = 1 or cggh/cgghh = 0; combinations such as ctt ~ 0.2 with chhh ~ 2.5 and cgghh ~ 0.3 are not examined and could alter the cluster-to-shape associations and the benchmark points.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript classifies shapes of the NLO gg->HH invariant mass distribution computed with full top-quark mass dependence in a five-coupling HEFT parameterization. It first defines four pre-defined shape types and maps them onto ten 2D coupling slices, then applies an autoencoder plus KMeans clustering to identify four or seven shape clusters and derives seven benchmark points from the cluster centers. The paper claims that the unsupervised method captures subtle shape features, such as enhanced tails and shoulders, better than the predefined taxonomy, and concludes that small deviations of ctt from zero tend to produce doubly peaked mhh distributions while chhh drives the low-mass enhancement.","tokens_in":18214,"tokens_out":5329,"duration_ms":49825,"significance":"If the claims hold, the seven-cluster taxonomy and the NLO benchmark points are a useful contribution: they extend the LO cluster-analysis benchmarks of Ref. [62] to full NLO with top-mass dependence, and they demonstrate a transparent unsupervised pipeline based on public coefficient tables. The classification logic is simple and reproducible, and the paper is honest about the arbitrariness of the pre-defined shapes. However, the mapping to the five-dimensional parameter space and the 'very well' performance claim are not yet quantitatively established because of the 2D-slice sampling and the absence of cluster-quality metrics.","major_comments":[{"comment":"The paper's central claim that the analysis maps shape classes onto the five-dimensional coupling space is not supported by the sampling described. Section 2.3 states that all two-coupling scans set the remaining three couplings to their SM values, and the cluster maps in Section 3.2 are drawn on the same ten 2D planes. Because Eq. (2.5) contains terms involving three or more couplings (e.g., A9 ctt cggh chhh, A17 ct ctt cggh, A18 ct cggh^2 chhh), simultaneous deviations can produce mhh shapes that never appear in any 2D slice with the remaining couplings SM-valued. The benchmark selection in Section 3.3 searches the same input grids, which makes the multiple non-SM benchmark points in Table 2 (e.g., point 1 with ct=0.94, chhh=3.94, ctt=-1/3, cggh=0.5, cgghh=1/3 as rendered in the manuscript) difficult to reconcile with a purely 2D-slice input set. The authors should either generate and scan a genuine 5D grid (or a structured sampling of the full space) and re-derive the cluster maps and benchmarks, or explicitly restrict the scope of the conclusions to the 2D-slice families.","section":"§2.3, §3.2, §3.3"},{"comment":"The choice of seven clusters and the claim that the unsupervised procedure 'captures shape features very well' are not quantitatively validated. Section 3.1 reports that seven clusters 'seemed to be the optimal number' based on visual inspection of the cluster centers, and the comparison in Section 3.2 is made against the authors' own four predefined shapes. No objective cluster-quality metric (e.g., silhouette score, Davies-Bouldin index, reconstruction error as a function of latent dimension, or stability across the ten encoder models) is provided. Since the central claim is the superiority of the unsupervised taxonomy, this omission is load-bearing; a quantitative validation would also make the seven-cluster choice reproducible.","section":"§3.1"}],"minor_comments":[{"comment":"There are several typos: 'deﬁnine' in Section 1, 'gobal' in Section 3.1, and 'disribution' in Section 3.2.","section":"§1, §3.1, §3.2"},{"comment":"The text and the caption refer to 'Table 3.3' when the benchmark points are in Table 2; please fix the cross-reference.","section":"§3.3 and Fig. 19 caption"},{"comment":"The term 'A10cttccgghh' appears to be a typesetting error for A10 ctt cgghh; please correct it.","section":"Eq. (2.5)"},{"comment":"Figure 12 is never discussed in the body of the paper; either refer to it explicitly in Section 3.2 or remove it.","section":"Fig. 12"},{"comment":"The benchmark table is difficult to read because the fractional entries are split across lines; please use consistent decimal or fraction notation for all couplings.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of JHEP. The main gap is the mismatch between the 2D-slice sampling and the claimed five-dimensional parameter-space mapping; this is fixable by extending the scan or softening the claims. I do not see a correctness error in the physics input, and the paper is not circular in any problematic sense."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Capozi and Heinrich do something genuinely useful: they take the established full-NLO, full-mt di-Higgs calculation and turn the mhh shape variation into a small taxonomy of seven clusters with concrete benchmark points. The NLO input is solid (Buchalla et al. 2018), the sampling of 10^5 distributions is reasonable, and the benchmark selection is transparent. The main new deliverable—seven NLO benchmark points plus the cluster maps—is a real resource for experimental interpretation. The novelty claim is fair: previous cluster analyses were LO and used predefined shapes; the autoencoder plus KMeans step is not exotic, but combining it with full-NLO shapes is new.\n\nThe soft spots are mostly about validation, not about the physics input. First, the claim that the unsupervised method captures shape features very well is supported only by eye: there is no reconstruction error, silhouette score, or stability metric beyond the cluster-center plots. Second, Section 2.2's shape types are subjective, and the paper says so; that is acceptable if the unsupervised part is meant to remove that bias, but the evaluation then compares the ML clusters to the same human-defined types. Third, the 20% exclusion for type 4 points under statistical uncertainties is a real selection effect; it should be reported in the benchmark discussion, not buried in Section 2.2.\n\nThe stress-test concern about 2D slices is legitimate. Section 2.3 fixes the other three couplings to SM in every projection, and the cluster maps and benchmark-point search inherit that. Eq. (2.5) has genuine mixed triple-coupling terms such as ctt cggh chhh and ct cggh^2 chhh, so simultaneous deviations can produce shapes that no 2D slice sees. The paper does not sample the full 5D space, so the maps are more like a set of instructive slices than a global shape atlas. This is the main limitation, and it should be stated explicitly.\n\nI would send the paper to a serious referee. It is not a breakthrough, but it is a clean phenomenological resource with reproducible underlying input, and the benchmark points will be cited by LHC analysers. I would ask for code/data release and, ideally, one full 5D check or a clear statement that the maps are slice-based only.","headline":"A useful NLO shape taxonomy with honest limits; the 2D-slice sampling is the main caveat, not the ML.","tokens_in":18804,"tokens_out":2593,"would_cite":true,"duration_ms":27546,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Unsupervised learning resolves seven shape classes in Higgs pair mass spectra that hand-defined types blur.","keywords":["Higgs pair production","shape analysis","anomalous couplings","non-linear EFT","autoencoder","KMeans clustering","benchmark points","next-to-leading order QCD"],"falsifier":"Evaluate $m_{hh}$ shapes at parameter points where $c_{hhh}$, $c_{tt}$, and $c_{gghh}$ are all shifted from their SM values, using the published coefficient tables; if such points produce shapes outside the regions predicted by the two-coupling maps, the maps are not representative of the full five-dimensional space.","tokens_in":17715,"feed_emoji":"⚛️","tokens_out":6299,"duration_ms":66150,"temperature":0.7,"pith_summary":"This paper tries to establish that the shape of the Higgs boson pair invariant mass distribution is a sharp diagnostic of anomalous Higgs couplings, and that an unsupervised classifier extracts that information more faithfully than a small set of human-defined shapes. Working at NLO with full top-quark mass dependence in a five-coupling non-linear EFT, it maps which regions of coupling space produce which $m_{hh}$ shapes. It claims the unsupervised approach resolves features such as an enhanced tail, a shoulder, and close-by double peaks that predefined categories blur, and yields seven cluster centers usable as benchmark points. If right, experimental searches can use these shapes and benchmark points to constrain couplings, especially $c_{hhh}$ and $c_{tt}$, beyond what total cross-section limits alone provide.","feed_headline":"Seven shape clusters map anomalous Higgs couplings at NLO","feed_subtitle":"Machine learning on Higgs-pair mass spectra finds seven reproducible shapes and seven benchmark points.","key_machinery":"The load-bearing object is the factorization of the NLO differential cross section into a sum over coupling monomials times bin-wise coefficient functions $A_i$ (Eq. 2.5), which allows a dense scan of the five-dimensional coupling space. On top of this sits an autoencoder that compresses each normalised 30-bin $m_{hh}$ histogram into a four-dimensional latent vector, followed by a standard K-means clustering algorithm into a chosen number of shape clusters. Ten differently initialised encoder models are trained, and a majority vote assigns each parameter point its final cluster label. The cluster centers act as shape prototypes, and a distance-based procedure converts them into concrete benchmark points in coupling space.","core_discovery":"The central claim is that an autoencoder followed by K-means clustering on normalised NLO $m_{hh}$ distributions identifies seven reproducible shape classes whose cluster centers are stable across ten encoder models, and that the parameter-space maps built from these clusters are more discriminating than the four predefined shape types. In particular, the paper claims that small deviations of $c_{tt}$ from zero are very likely to produce a doubly peaked $m_{hh}$ structure, while SM-like shapes reappear as $c_{tt}$ moves further away from zero. It also claims that shapes with an enhanced tail or a shoulder are likely to be produced by nonzero values of $c_{gghh}$, and that shape variation is dominated by $c_{hhh}$ and $c_{tt}$. The paper derives seven NLO benchmark points from the cluster centers, each satisfying the current combined LHC upper bound of 6.9 times the Standard Model cross section.","pith_inferences":["A natural extension is to treat the cluster centers as a continuous latent space and interpolate between them, which could give smooth parameter-shape maps rather than discrete labels; the paper only provides discrete clusters and benchmark points.","Since the coefficient tables are published for 13, 14, and 27 TeV, the same clustering pipeline could be rerun at 14 and 27 TeV; the paper's benchmark points are quoted only for 13 TeV.","If small $c_{tt}$ deviations really do imprint a double peak, that shape may be one of the cleanest new-physics signatures in di-Higgs data, provided background and parton-shower modeling retain the feature.","The autoencoder's latent dimension of four limits the resolution of the shape manifold; a larger latent space might separate additional features such as the exact peak separation, which could be tested before applying the method to data."],"forward_implications":["The seven cluster centers give experimentalists concrete NLO benchmark points, each within the current combined LHC cross-section limit, for profile-likelihood or template fits.","The claim that small $c_{tt}$ deviations produce a doubly peaked $m_{hh}$ structure offers a direct target: search for that shape to constrain $c_{tt}$, which single-Higgs measurements constrain only weakly.","Shapes with an enhanced tail or a shoulder are tied to nonzero $c_{gghh}$, so shape analyses can probe the effective gluon-Higgs couplings that total cross sections alone do not resolve.","Because $m_{hh}$ is more shape-sensitive than $p_{T,h}$, differential $m_{hh}$ measurements should be prioritized in future di-Higgs analyses; the method itself transfers to other observables and other processes.","In the SMEFT limit, where $c_{ggh}$ and $c_{gghh}$ are related and $c_{tt}$ is suppressed relative to $c_t$, the shape maps reduce to a three-dimensional parameter space that can be visualised directly and used for more model-dependent projections."],"supporting_citations":[{"why":"Supplies the NLO differential cross-section coefficient functions $A_i$ (Eq. 2.5) with full top-quark mass dependence used to generate all $m_{hh}$ shapes.","marker":"[71]"},{"why":"The earlier LO cluster-analysis benchmark definition whose bias and degeneracy this paper aims to overcome.","marker":"[62]"},{"why":"Experimental combined search limit on the gluon-fusion di-Higgs cross section, used to exclude parameter regions and select benchmark points.","marker":"[1]"},{"why":"The complementary LHC limit shown in the parameter-space maps, bounding the allowed coupling regions.","marker":"[2]"},{"why":"Provides the K-means clustering algorithm used to group the autoencoded shapes.","marker":"[74]"},{"why":"Provides the autoencoder implementation used for the unsupervised shape compression.","marker":"[107]"},{"why":"Backend used to train the autoencoder models.","marker":"[108]"}],"fun_headline_variants":["ML identifies 7 reproducible Higgs-pair shape classes","Autoencoder + K-means yield 7 Higgs shape clusters","Unsupervised learning maps Higgs-pair shapes to couplings","Seven robust shape clusters from NLO Higgs-pair spectra","Doubly peaked Higgs shapes signal small c_tt deviations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's maps and conclusions assume that two-dimensional slices with the other three couplings at their Standard Model values represent the full five-dimensional coupling space, so shapes caused by simultaneous deviations of three or more couplings could be missed.","fun_headline_variants_meta":{"raw":{"variants":["ML identifies 7 reproducible Higgs-pair shape classes","Autoencoder + K-means yield 7 Higgs shape clusters","Unsupervised learning maps Higgs-pair shapes to couplings","Seven robust shape clusters from NLO Higgs-pair spectra","Doubly peaked Higgs shapes signal small c_tt deviations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000713,"raw_usage":{"total_tokens":3153,"prompt_tokens":834,"completion_tokens":2319,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":450,"completion_tokens_details":{"reasoning_tokens":2238}},"tokens_in":450,"tokens_out":2319,"duration_ms":17613,"temperature":1.0,"reasoning_tokens":2238,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:24:56.343034+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Evaluate $m_{hh}$ shapes at parameter points where $c_{hhh}$, $c_{tt}$, and $c_{gghh}$ are all shifted from their SM values, using the published coefficient tables; if such points produce shapes outside the regions predicted by the two-coupling maps, the maps are not representative of the full five-dimensional space.","supporting_citations":[],"review_version":1}