{"id":"167d7776-b6d2-4d5f-91b5-302b77ee3e0c","arxiv_id":"2507.17884","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A CMS reinterpretation of the 138 inverse femtobarn dijet-pair search finds the 8.6 TeV excess persists when the mediator width is increased to 10%, with local significance 3.6 standard deviations.","lead":"CMS re-analyzed its earlier search for pairs of dijet resonances, now allowing the intermediate particle to be broad (natural width from 1.5% to 10% of its mass). The excess near a four-jet mass of 8.6 TeV keeps a local significance above 3.6 standard deviations, so a broad resonance is presented as an equally valid interpretation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Broad-resonance stability at 8.6 TeV rests on a simulated low-mass signal tail (40% below 5.8 TeV, Sec. 4) that is neither independently validated nor assigned a specific systematic, so the width-insensitivity claim may be an artifact of the generator/pairing model.","rationale":"The reader correctly identified the simulated broad-signal shape, especially the low-mass tail containing the 5.8 TeV candidate, as the weakest assumption. My reading confirms this is the load-bearing point: the paper's own numbers show the local significance falling from 3.9 to 3.6 s.d. as the width increases, and the 10% template maintains significance only by adding a large low-mass tail that is not independently validated. The paper provides no comparison between the simulated mispairing rate and any data control region, and the inherited systematic uncertainties do not cover this component. The overstatement in the abstract about 'equally valid' is secondary; the quantitative stability claim itself is what needs scrutiny. Because the concern is material but addressable with a targeted refit, the CONDITIONAL verdict from the reader remains appropriate, so no adjustment is needed.","tokens_in":39467,"tokens_out":8588,"duration_ms":97353,"concrete_test":"Refit the Gamma/M=10% row of Table 2 at MY=8.6 TeV after modifying the signal template with a nuisance parameter that multiplies only the m4j < 6.5 TeV component, profiled against data. If the fitted weight is far from unity or the local significance drops below 3 s.d., the width-stability conclusion is not robust; if the significance stays near 3.6 s.d. for weights spanning 0.5x to 2x, the concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the 8.6 TeV excess is equally compatible with broad resonances depends on the 10%-width signal template placing about 40% of its yield at m4j <= 5.8 TeV (Sec. 4), exactly the region of the third candidate event. This low-mass tail is not a feature of the narrow search being reinterpreted; Sec. 3 (Fig. 2) attributes the tail-like structure in the 10% contour to incorrect jet combinations assigned by the pairing algorithm. No closure test, control sample, or generator-level comparison is presented to validate the rate and shape of this tail, and the systematic uncertainties listed in Sec. 5 (jet energy scale, mass resolution, luminosity, background) do not include a dedicated uncertainty for the mispairing/radiation tail. Because the tail dominates the broad-signal likelihood and is what incorporates the 5.8 TeV event, an unquantified error in that component would change the quoted 3.6 s.d. and undermine the 'equally valid' conclusion. The analysis itself is not internally inconsistent; rather, its headline robustness claim is conditional on an unvalidated part of the simulation chain.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reinterprets the CMS search for pairs of dijet resonances (Ref. [1]) by allowing the intermediate mediator Y to have a natural width up to 10% of its mass, using the same 138 fb^-1 of 13 TeV proton-proton collision data. The analysis reuses the event selection, jet pairing, and background-fitting procedure of Ref. [1], with signal templates generated for a diquark benchmark model at widths of 1.5%, 5%, and 10%, for four-jet resonance masses from 2 to 10 TeV and for several dijet-to-four-jet mass ratios. Upper limits at 95% CL on the signal cross section and local/global significances are obtained from a multibin Poisson likelihood in 13 bins of alpha = m2j/m4j. The central result is that the excess at a four-jet mass of 8.6 TeV retains a local significance of 3.9 to 3.6 s.d. as the width increases from 1.5% to 10%, leading the authors to conclude that a broad resonance is an equally valid interpretation of the excess. A second excess at 3.6 TeV is also reported, with local (global) significance up to 3.9 (2.2) s.d. at 10% width.","tokens_in":39661,"tokens_out":8073,"duration_ms":81238,"significance":"If the results are correct, this is a valuable reinterpretation: it demonstrates that the previously reported 8.6 TeV excess is not an artifact of the narrow-width assumption and that broad mediators remain a viable explanation. The analysis uses the established CMS statistical framework, including a multibin Poisson likelihood, background-only fits with reported p-values, discrete profiling for background functional-form uncertainty, and pseudo-experiment cross-checks for small event counts. Results are provided in the HEPData record, which is commendable for reproducibility. The paper extends the scope of the original search by scanning over widths and providing model-independent limits in the plane of four-jet mass and dijet mass ratio. The main scientific interest is the claimed width-insensitivity of the excess, which rests on a simulated low-mass signal tail; this is also the main technical risk, as discussed below.","major_comments":[{"comment":"The central claim that the 8.6 TeV excess is equally compatible with broad resonances relies on the 10%-width signal template placing about 40% of its yield at m4j <= 5.8 TeV (Section 4), a region that contains the third candidate event. The text attributes this low-mass tail to incorrect jet combinations assigned by the pairing algorithm (Section 3, Fig. 2), but no closure test, control-sample validation, or generator-level comparison is presented to establish the rate and shape of this tail, and the list of systematic uncertainties in Section 5 does not include a dedicated uncertainty for this mispairing component. Because the tail dominates the broad-signal likelihood and is what allows the 5.8 TeV event to support the broad hypothesis, the reported width-insensitivity of the significance (3.9 to 3.6 s.d.) is not established without quantifying this component. The authors should validate the mispairing tail with, for example, a data control region or an alternate jet-pairing algorithm, or assign and propagate a corresponding systematic uncertainty.","section":"Section 4, Fig. 3; Section 3, Fig. 2"},{"comment":"The mass limits in Table 1 and the model comparisons in Figs. 8-10 use a cross-section prediction corrected by an 'efficiency factor that isolates the m4j range where the expected significance exceeds 99% of that obtained without the correction.' This definition is self-referential because the theoretical prediction used for the exclusion is corrected using the significance of the signal being tested, and the exact m4j range and the significance computation are not specified. As a result, the reported exclusion limits (for example, 8.8 TeV for Suu at 10% width) depend on a criterion that is not fully reproducible. The authors should provide a precise mathematical definition of this factor and demonstrate that the excluded mass regions are stable under reasonable variations of the efficiency threshold.","section":"Section 5; Table 1"}],"minor_comments":[{"comment":"The global significance is quoted separately for each width (1.6, 1.5, and 1.4 s.d. in Table 2), but it is not stated whether the trials factor includes the discrete scan over the three width hypotheses. If not, the reported global p-values are conditional on the width and should be interpreted with that caveat.","section":"Section 5"},{"comment":"The figure labels and caption use 'SM 68% contours' for the simulated signal contours; this is ambiguous because 'SM' usually denotes the standard model. Consider renaming them to 'signal 68% contours' for clarity.","section":"Section 3, Fig. 2"},{"comment":"The statement that 'diquark signals with a width larger than 10% do not exhibit a peak at the resonance mass' is presented without a reference or illustration. A supporting figure or quantitative criterion would make this modeling choice easier to assess.","section":"Section 4"},{"comment":"The vector-like quark width is reported to reach up to 3.7% at a mass of 4.2 TeV, but the analysis assumes that X is a narrow dijet resonance. Since the search targets narrow X, it would be useful to state explicitly that a 3.7% width is negligible compared to the experimental resolution or to restrict the affected mass range.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is essentially a reinterpretation of a published CMS search, and the novelty lies in the width scan of the 8.6 TeV excess. The statistical framework is sound and follows CMS standards, but the central robustness claim rests on an unvalidated simulated mispairing tail. I recommend major revision, requesting either a validation of that tail or a dedicated systematic uncertainty for it. The efficiency correction for model exclusion limits is also not fully specified and should be clarified. These points are addressable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a straightforward reinterpretation of CMS's earlier narrow four-jet resonance search, not a new measurement. The genuinely new pieces are the signal templates for 1.5, 5, and 10% widths, the 95% CL limits as functions of width, and the explicit demonstration that the 8.6 TeV excess stays at 3.6–3.9 s.d. local across widths. That is useful: it tells you the excess is not an artifact of the narrow-width assumption. The statistical framework is the same well-tested procedure from Ref. [1], with data-driven background fits in thirteen alpha bins, discrete profiling, Combine, and a HEPData record. No circularity: widths are scanned inputs, background is fitted to data, significances come from the observed likelihood. The 3.6 TeV second excess at 10% width (global 2.2 s.d.) is also worth noting, and Section 6's presentation of the third candidate event is clear.\n\nThe main soft spot is wording. The abstract says a broad resonance is an \"equally valid interpretation\" of the excess, but the local significance actually decreases from 3.9 to 3.6 s.d. as the width increases, and no formal narrow-versus-broad hypothesis test is provided. \"Compatible\" is the defensible claim, not \"equally valid.\"\n\nThe bigger concern is the one the stress-test note flags. For a 10% width, 40% of the signal is expected at m4j below 5.8 TeV, and that low-mass tail is what pulls in the third candidate event. The paper itself attributes the tail to incorrect jet combinations from the pairing algorithm. That simulation component is load-bearing for the width-insensitivity claim, and it gets no dedicated systematic uncertainty or closure test. If the mispairing rate or shape is wrong, the broad-significance numbers could move. This is not fatal—it is a simulation-modeling uncertainty in an otherwise standard analysis—but it should be addressed with a control sample, generator-level comparison, or an explicit nuisance parameter.\n\nOverall, the math and data handling look solid, and the citation pattern is appropriate; this is a legitimate extension of Ref. [1], not an overreach. The paper deserves serious peer review. I would send it to referees with a request for validation of the mispairing tail and a toned-down abstract claim. It is not a discovery paper and does not claim to be one, but as a limits-and-interpretation paper it is useful and honest.","headline":"A solid, honest reinterpretation that shows the 8.6 TeV excess persists across widths, but the 'equally valid' claim overstates and the 10% result leans on an unvalidated simulated low-mass tail.","tokens_in":40287,"tokens_out":2285,"would_cite":true,"duration_ms":27248,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The 8.6 TeV four-jet excess remains significant for resonance widths up to 10 percent, so a broad mediator is as valid as a narrow one.","keywords":["broad mediator","dijet resonance pair production","four-jet final state","resonance width","diquark model","8.6 TeV excess","LHC proton-proton collisions","13 TeV"],"falsifier":"Compare the number and distribution of events with $m_{4j} \\le 5.8$ TeV and $m_{2j}\\approx 2$ TeV in the full Run 2 dataset against the 40% low-mass-tail prediction for the 10% width hypothesis, or rerun the analysis with a different jet-pairing algorithm and check whether the 8.6 TeV significance still stays above 3.6 standard deviations at 10% width.","tokens_in":39173,"feed_emoji":"⚛️","tokens_out":8781,"duration_ms":82101,"temperature":0.7,"pith_summary":"This paper reinterprets an earlier search for pairs of dijet resonances, allowing the parent four-jet resonance to be broad rather than narrow. Using the full $138~\\text{fb}^{-1}$ dataset from $\\sqrt{s}=13$ TeV proton-proton collisions, it shows that resonances with natural widths of 1.5, 5, and 10% of their mass all describe the observed excess near a four-jet mass of 8.6 TeV, with local significance 3.9 to 3.6 standard deviations and global significance 1.6 to 1.4 standard deviations. The key addition is that a 10% width naturally accommodates a third candidate event at 5.8 TeV that narrow-resonance fits barely use. This matters because it means the reported excess does not depend on the narrow-width assumption, and it provides model-independent limits on heavy resonances decaying to two equal-mass dijet pairs.","feed_headline":"The 8.6 TeV excess survives widths up to 10 percent","feed_subtitle":"A 5.8 TeV event that narrow fits ignore supports a broad parent resonance.","key_machinery":"The search is organized around the dimensionless ratio $\\alpha = m_{2j}/m_{4j}$, the average dijet mass over the four-jet mass, which separates signal from the smoothly falling QCD background. Data are split into thirteen $\\alpha$ bins, and in each bin the four-jet mass distribution is fitted with a three-parameter modified dijet background function, with alternative functions used as discrete nuisances. Signal templates are generated for the diquark model with widths 1.5, 5, and 10% by tuning the couplings $y_{uu}$ and $y_\\chi$ in Eqs. (3)--(4); the broad templates develop a long low-mass tail that arises from incorrect jet combinations assigned by the pairing algorithm. A multibin Poisson likelihood, profiled over background and systematic nuisance parameters, yields the significances and 95% CL limits.","core_discovery":"The central claim is that a broad mediator, with width up to 10% of its mass, is an equally valid interpretation of the 8.6 TeV four-jet excess previously reported under a narrow-resonance assumption. For a diquark benchmark with mass ratio $\\alpha_{\\text{true}} = M_X/M_Y = 0.25$, the local significance at $M_Y = 8.6$ TeV remains 3.9, 3.8, and 3.6 standard deviations for widths 1.5, 5, and 10%, while the global significance stays between 1.6 and 1.4 standard deviations. The broader signals extend to lower four-jet masses through events in which the jet-pairing algorithm combines the wrong jets; for a 10% width, 40% of the signal is expected at $m_{4j} \\le 5.8$ TeV. This tail is what makes the event at $m_{4j}=5.8$ TeV, $m_{2j}=2.0$ TeV compatible with a broad resonance, supporting the interpretation. The analysis also reports a separate excess at $M_Y=3.6$ TeV, $\\alpha_{\\text{true}}=0.29$, whose local (global) significance rises to 3.9 (2.2) standard deviations at 10% width, and presents 95% CL upper limits on $\\sigma B A$ for masses between 2 and 10 TeV.","pith_inferences":["If the broad-mediator interpretation is correct, the same 8.6 TeV excess should reappear in the upcoming higher-statistics Run 3 data, with the ratio of tail to peak events following the width-dependent prediction.","A combined fit of the CMS and the other experiment's candidate events in the two-dimensional $(m_{4j}, m_{2j})$ plane, rather than separate $\\alpha$ bins, could sharpen the broad-versus-narrow discrimination.","The mis-pairing mechanism that creates the low-mass tail is testable with generator-level studies: if a different pairing or a boosted reconstruction removes the 5.8 TeV event from the signal region, the broad-resonance interpretation loses its main supporting event.","The 3.6 TeV excess's growth with width suggests it may be the same underlying component as the nonresonant dijet-pair effect seen in prior data; a dedicated search with finer $\\alpha$ bins could clarify whether it is resonant at all."],"forward_implications":["The 8.6 TeV excess can no longer be attributed solely to a narrow-width assumption; searches with broad mediators must be included in its interpretation.","The event at $m_{4j}=5.8$ TeV, which a narrow fit barely uses, becomes a signal-like member of the broad-resonance hypothesis and improves the fit.","The independent candidate event at $m_{4j}=6.6$ TeV and $m_{2j}=2.2$ TeV falls inside the 5% and 10% width contours, so the two experiments' observations are mutually consistent under a broad mediator.","For the diquark benchmark with $\\alpha_{\\text{true}}=0.25$ and 10% width, the $S_{uu}$ model is excluded at 95% CL above 8.8 TeV while $S_{dd}$ remains viable at 8.6 TeV, so the excess can still be a scalar diquark hypothesis.","The second excess at 3.6 TeV, which grows with width, provides an independent target for future searches and for model building."],"supporting_citations":[{"why":"Supplies the dataset, event selection, jet-pairing algorithm, background-fitting procedure, and the narrow-resonance excess that this paper reinterprets.","marker":"[1]"},{"why":"Provides the independent candidate event at a four-jet mass of 6.6 TeV and average dijet mass of 2.2 TeV that falls inside the 5% and 10% width contours.","marker":"[2]"},{"why":"Defines the scalar diquark benchmark model, with couplings and partial-width formulas used to generate the 1.5, 5, and 10% resonance widths.","marker":"[3]"}],"fun_headline_variants":["Broad resonance equally valid at 8.6 TeV, width to 10%","Significance persists to 10% width for 8.6 TeV excess","Event at 5.8 TeV backs broad parent at 8.6 TeV","Two excesses, 8.6 and 3.6 TeV, fit broad widths"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The weakest link is the simulation's prediction of the broad-resonance signal shape, especially the low-mass tail produced by mis-paired jets; if the rate or shape of those tail events is miscalibrated, the claimed stability of the significance with width would not hold.","fun_headline_variants_meta":{"raw":{"variants":["Broad resonance equally valid at 8.6 TeV, width to 10%","Significance persists to 10% width for 8.6 TeV excess","Event at 5.8 TeV backs broad parent at 8.6 TeV","Two excesses, 8.6 and 3.6 TeV, fit broad widths"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000451,"raw_usage":{"total_tokens":2416,"prompt_tokens":1237,"completion_tokens":1179,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":853,"completion_tokens_details":{"reasoning_tokens":1086}},"tokens_in":853,"tokens_out":1179,"duration_ms":9962,"temperature":1.0,"reasoning_tokens":1086,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:17:31.032127+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the number and distribution of events with $m_{4j} \\le 5.8$ TeV and $m_{2j}\\approx 2$ TeV in the full Run 2 dataset against the 40% low-mass-tail prediction for the 10% width hypothesis, or rerun the analysis with a different jet-pairing algorithm and check whether the 8.6 TeV significance still stays above 3.6 standard deviations at 10% width.","supporting_citations":[],"review_version":1}