{"id":"6bb1fd5e-b887-4c41-89c3-7cc65ba4025a","arxiv_id":"2412.13976","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An independent validation of the ATLAS EXOT-2019-23 reinterpretation map shows it works well for high-pT benchmarks but fails for low-mass and short-lifetime signals.","lead":"This note tests the efficiency map ATLAS released for recasting its search for long-lived particles decaying into displaced hadronic jets. It shows the map reproduces the published limits well for high momentum benchmarks but is less reliable at low momentum and short lifetimes.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The recasting validation's main load-bearing assumption is that MG+Pythia8 truth kinematics match ATLAS's internal simulation; a cross-generator test would settle whether the high-ET agreement is robust or coincidental.","rationale":"The reader's weakest_assumption identifies the same core issue: the validation relies on the faithfulness of the external event generation to ATLAS's internal simulation, without an independent test. This is indeed the most load-bearing assumption because the entire validation procedure in Section IV compares recast outputs to published ATLAS results, and any mismatch in generator-level kinematics would propagate through the map into the efficiency and limit curves. However, the paper carefully limits its claim to the specific implementation using MG+Pythia8 and to high-ET benchmarks, and it explicitly documents several related limitations (e.g., eta-distribution dependence, region-A-only approximation, trackless and missing-energy assumptions, and lifetime-dependent validity). The observed agreement for the high-ET benchmarks is a nontrivial consistency check that supports the claim as stated. The absence of a cross-generator test is a legitimate residual risk, but it does not invalidate the conclusion because the conclusion is not a universal claim about all generators or all models. The proposed Herwig test would provide a decisive robustness check and could be added as a future improvement, but the current verdict of ACCEPT remains appropriate given the paper's transparency and the released code. The correctness risk is already marked medium, and the concern we raise is within that risk. Therefore, we do not adjust the verdict, but we flag the test as a worthwhile verification step.","tokens_in":6956,"tokens_out":9639,"duration_ms":90883,"concrete_test":"Regenerate the high-ET benchmark points (m_phi = 1000, 600, 400 GeV; m_s = 275, 150, 100 GeV) using the same MadGraph5_aMC@NLO hard process and HAHM model, but shower them with Herwig 7 instead of Pythia8. Apply the published High-ET efficiency map exactly as described in Section III-B, computing the Bin Index and summing per-event probabilities to obtain recast efficiencies and limits. Compare these new results to the ATLAS published efficiency curves (Fig. 5) and to the Pythia8-based results. If the Herwig-based efficiencies deviate from the ATLAS values by more than the map's stated 25% uncertainty for efficiencies above 0.5%, the validation is generator-dependent and the agreement with Pythia8 could be coincidental. If the results remain within uncertainties, the concern is settled and the recasting procedure is robust to showering choices, at least for high-ET benchmarks.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the recasting procedure approximates the ATLAS results for high-ET benchmarks rests on the assumption, introduced in Section III-B, that events generated with MadGraph5_aMC@NLO and showered with Pythia8 at default settings reproduce the truth-level LLP kinematics (transverse momentum, decay position, decay products) that enter the efficiency map with sufficient fidelity to ATLAS's internal Monte Carlo. The validation in Section IV compares only against the same HAHM benchmark model used to derive the map, so the agreement shown in Figs. 5 and 7 could reflect internal consistency of the map rather than robustness to generator choices. If Pythia8's parton shower and hadronisation shift the LLP pT distribution (e.g., via ISR recoil) relative to ATLAS's tuned generator, the map would apply incorrect per-bin probabilities, and the observed agreement for high-ET benchmarks would be coincidental. The paper demonstrates the effect of hadronisation on jet pT (Figs. 3 and 4) but does not test an alternative showering algorithm or tune, nor does it quantify how much the LLP pT distribution itself changes. Since the map is binned coarsely (5 pT bins), even a modest shift across bin boundaries could bias efficiency estimates. The paper does list this as an implicit assumption and calls out other model-dependence limitations, but no independent check is provided. The claim as stated—'managed to approximate the published results in a satisfactory manner'—is established for the specific setup used, but the generality of that conclusion for external users with different generators remains untested. A dedicated cross-generator test would directly probe this load-bearing assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This note validates the ATLAS-EXOT-2019-23 reinterpretation efficiency map, which maps truth-level LLP kinematics (decay position, transverse momentum, decay-product PDGIDs) to a signal-region selection probability. The authors generate HAHM events with MadGraph5_aMC@NLO and Pythia8 for six benchmark points, apply the published map, and compare the resulting efficiency curves and cross-section limits with the ATLAS publication. They find good agreement for high-ET benchmarks, degraded agreement for low-ET benchmarks, and they catalogue practical obstacles facing external users of the ATLAS material. The paper also provides open-source code implementing the recasting procedure.","tokens_in":7254,"tokens_out":4737,"duration_ms":46384,"significance":"If the map is reliable, it is a valuable and computationally cheap tool for reinterpretation of an important displaced-jet search, and this note is a useful independent test of that tool. The authors are candid about the limitations they find, and the code release is a concrete contribution that others can reuse. The validation is, however, a closure test against the same ATLAS results that motivated the map, using a single generator setup and a small number of benchmark points; the agreement shown is therefore not a fully independent check of the map's generality. The paper's central claim, that the procedure approximates the published results satisfactorily for high-ET benchmarks, is nevertheless supported by the comparisons shown.","major_comments":[],"minor_comments":[{"comment":"The limit curves in Figs. 7 and 8 are presented without propagating the map uncertainties quoted in Section II-B (25% for High-ET, 33% for Low-ET). Since the map accuracy is limited, the visual agreement in Fig. 7 may be partly fortuitous. Please either add uncertainty bands to the recast limits or explicitly state in the text and captions that these are central values only, so that readers can judge the significance of the agreement.","section":"Section IV-B, Figs. 7-8"},{"comment":"The validation uses only MadGraph5_aMC@NLO with Pythia8 at default settings and only the HAHM benchmark used by ATLAS. The paper should state more explicitly that the demonstrated agreement is a closure test for this particular generator setup, and that robustness to alternative shower tunes or generators has not been tested. This would prevent readers from over-generalizing the 'satisfactory' verdict beyond the tested conditions.","section":"Section III-B and Section VI"},{"comment":"In the bullet about hadronisation effects, 'Fig, 3' should be 'Fig. 3'.","section":"Section V-A"},{"comment":"The caption reads 'm_s = 100, 55 GeV', but Table I lists m_s = 50 GeV for m_phi = 200 GeV and m_s = 55 GeV for m_phi = 125 GeV. Please correct the inconsistency or clarify the intended values.","section":"Figure 8 caption"},{"comment":"The sentence 'As we can see in Fig. 2, the map is symmetric between the two LLPs, The choice of LLP 1 and 2 is arbitrary' contains a grammatical error; consider rewording to 'As we can see in Fig. 2, the map is symmetric between the two LLPs, so the choice of LLP 1 and 2 is arbitrary.'","section":"Section II-A"},{"comment":"'Nbr events' should be spelled out as 'Number of events' in the caption.","section":"Table I"}],"recommendation":"minor_revision","confidential_remarks":"This is a technical note rather than a discovery paper, but it provides a useful service to the recasting community by testing public reinterpretation material and releasing code. The main limitation is that the validation is a closure test; the authors should be asked to make that scope explicit, but I do not see a need for additional external validation before publication in a suitable venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a solid, honest recasting note. The new content is the validation itself: six benchmark points, public code, and specific findings about the map's domain of validity. For high-pT benchmarks the map reproduces the ATLAS limits well; for low-pT and cτ below about 50 cm it doesn't, and hadronisation matters a lot at low mass. Those observations go beyond the original ATLAS paper and are genuinely useful for the LLP recasting community.\n\nWhat the paper does well: it ships code and data, which is the norm for recasting work and is done properly here. It is candid about the map's coarseness, the limited validation, and the ad-hoc map uncertainties. The list of obstacles for external users—inadequate MC documentation, missing model files, undocumented parameter choices—is a real contribution to the broader reproducibility discussion. The efficiency comparison figures are informative, and the physical explanation for why hadronisation is essential at low mediator mass (the pT threshold cannot be reached without shower recoil) is clear.\n\nSoft spots: the validation uses the same HAHM benchmark model that produced the map, with a single generator chain (MG+Pythia8 default). The high-pT agreement could partly reflect internal consistency of the map rather than robustness to generator choices. A cross-generator or cross-tune test would settle that, but the paper's central claim is carefully scoped: it says the procedure works 'in a satisfactory manner' for the setup used, and it flags limitations. I don't think this is fatal—it's a stated boundary, not a hidden one. The limit curves are shown without uncertainty propagation, which is a minor presentation issue. Six benchmark points are few, but they span a reasonable mass and lifetime range. The unchecked assumption that u,d,s,gluon decays behave like c-quarks is explicitly flagged, so readers are not misled.\n\nWho is this for? Phenomenologists doing LLP recasting, and anyone concerned with how experimental collaborations package auxiliary material for reinterpretation. It deserves serious peer review, not desk rejection. I would accept it and send to a knowledgeable referee, ideally one who has used the ATLAS map or built similar recasting tools. Recommended minor revisions: add a cross-generator test if feasible, or at least a paragraph on why it was not done, and propagate the map uncertainties to the limit curves.","headline":"Honest, useful validation of an ATLAS efficiency map for LLP recasting, with clear documentation of where it works and where it breaks.","tokens_in":7769,"tokens_out":1555,"would_cite":true,"duration_ms":15859,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper validates the efficiency map published with the EXOT-2019-23 search and shows that it reproduces the original limits for high-transverse-energy long-lived-particle models with lifetimes above about 50 cm.","keywords":["long-lived particles","displaced jets","recasting","efficiency map","Run-2 data","hadronisation","hidden Abelian Higgs model"],"falsifier":"Regenerate one of the high-ET benchmark points with a different shower tune or a different hadronisation model and push the events through the same map; if the resulting cross-section limits move by more than the quoted map uncertainties (about 25% for efficiencies above 0.5%), the validation's agreement is tune-dependent rather than robust. A second test: generate a model with a deliberately different pseudorapidity distribution and compare map-based limits with a full detector-level recast; if they disagree by more than a factor of a few, the map's folded acceptance assumption fails.","tokens_in":6798,"feed_emoji":"⚛️","tokens_out":6857,"duration_ms":58988,"temperature":0.7,"pith_summary":"This note asks whether the reinterpretation material released with the EXOT-2019-23 search can substitute for a full detector-level analysis when one wants to estimate a new model's signal — a procedure often called recasting. The material is a six-dimensional efficiency map that takes truth-level decay positions, transverse momenta, and decay-product identities for a pair of long-lived particles and returns the probability that the event would pass the search's signal selection. The authors regenerate the benchmark signal, push the generated events through the map, and compare the resulting selection efficiencies and cross-section limits with the published values. Their central finding is that the procedure recovers the published results reasonably well for high-transverse-energy benchmarks, is usable by an external analyst, but becomes less reliable for low-energy, low-mass benchmarks and for lifetimes below roughly 50 cm. The note also documents gaps in the public documentation that made the validation harder than it needed to be.","feed_headline":"Map reproduces displaced-jet limits for high-energy models","feed_subtitle":"Validated recasting tool is reliable above ~50 cm lifetimes; low-mass signals need care.","key_machinery":"The central object is the efficiency map: a binned lookup table, provided with the search, that maps truth-level long-lived-particle properties — decay position (transverse in the barrel, longitudinal in the endcap), transverse momentum, and decay-product type — to the probability that a pair of such decays would be selected in the signal region. It folds in trigger, reconstruction, machine-learning discriminants, and all analysis selections. The validation machinery works by generating benchmark events, computing each decay's bin index, reading per-event selection probabilities from the map, summing them into a sample efficiency, and converting that efficiency into a cross-section limit with a single-bin signal-plus-background fit using the published region-A yields.","core_discovery":"The paper's claim is that the published efficiency map is a serviceable recasting tool within a documented range, not a replacement for the full analysis everywhere. Concretely, for the high-ET selection the map-derived efficiencies and limits agree with the originally published curves once the map's stated uncertainties (about 25% for efficiencies above 0.5%) and validity limits are respected, with agreement best for large lifetime values. For the low-ET selection the procedure degrades: agreement is only to order of magnitude, hadronisation is essential, and one of the six benchmark points (mediator mass 60 GeV, LLP mass 5 GeV) yields efficiencies too low to be useful. The authors therefore recommend using the map only for lifetimes above about 50 cm, while noting that the map implicitly assumes new models resemble the training model in pseudorapidity distribution, tracklessness, and missing-hadronic-energy fraction.","pith_inferences":["A reasonable rule of thumb extending the paper's findings: do not quote a recast limit from this map for any model with mean proper lifetime below 50 cm, regardless of how well the high-lifetime part of the curve agrees.","The map's folded-acceptance assumption could be tested directly by generating a model with a deliberately different pseudorapidity distribution (for example, production via a vector-boson-fusion-like topology) and comparing map-based limits with a full-simulation recast; the paper identifies the assumption but does not quantify its failure.","If the same map format and validation procedure were applied to other searches, recasting could become a semi-automated check rather than a custom per-analysis exercise; the paper's wish list of standard formats points in that direction."],"forward_implications":["External users can reinterpret the search for new long-lived-particle models using only truth-level generator output, without rerunning detector simulation or the full analysis.","The recasting is reliable for high-transverse-energy benchmarks with mediator masses at or above a few hundred GeV, provided lifetimes exceed about 50 cm.","For low-mass, low-energy models, parton-shower and hadronisation effects are essential; omitting them pushes jet transverse momenta below threshold and badly distorts efficiencies.","The map silently assumes new models pass the trackless-jet and missing-hadronic-energy side requirements, so those assumptions must be checked before trusting a limit.","The public record currently lacks enough generator-level documentation to make independent validation straightforward, and improving that documentation is essential for recasting."],"supporting_citations":[{"why":"Supplies the published efficiencies, limits, and region-A yields that the validation procedure aims to reproduce.","marker":"[2]"},{"why":"Provides the efficiency map in structured form and the helper code that defines the bin indexing.","marker":"[3]"},{"why":"Defines the hidden-sector benchmark model used to generate the validation samples.","marker":"[5]"},{"why":"Companion reference for the hidden-sector model, grounding the mediator and dark-scalar setup.","marker":"[6]"},{"why":"Supplies the matrix-element generator used to produce hard-scattering events with the benchmark model.","marker":"[9]"},{"why":"Provides the showering and hadronisation stage that the paper shows is essential for low-energy agreement.","marker":"[10]"}],"fun_headline_variants":["Recasting map is reliable only above 50 cm lifetimes","Displaced-jet map reproduces high-ET limits but not low-ET","ATLAS efficiency map validated for high-ET, limited for low-ET","Recasting map works for high-ET, but low-ET is shaky","Map reliable for recasting only above 50 cm lifetimes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole validation rests on the assumption that truth-level kinematics from an external generator, after showering and hadronisation, match the generator-level distributions the original analysis used closely enough that the published efficiency map assigns the same selection probabilities; if the hadronisation model shifts jet momenta or decay positions differently from the original simulation, the agreement with the published limits could be a coincidence.","fun_headline_variants_meta":{"raw":{"variants":["Recasting map is reliable only above 50 cm lifetimes","Displaced-jet map reproduces high-ET limits but not low-ET","ATLAS efficiency map validated for high-ET, limited for low-ET","Recasting map works for high-ET, but low-ET is shaky","Map reliable for recasting only above 50 cm lifetimes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001226,"raw_usage":{"total_tokens":5021,"prompt_tokens":906,"completion_tokens":4115,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":522,"completion_tokens_details":{"reasoning_tokens":4021}},"tokens_in":522,"tokens_out":4115,"duration_ms":26544,"temperature":1.0,"reasoning_tokens":4021,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:35:29.414999+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Regenerate one of the high-ET benchmark points with a different shower tune or a different hadronisation model and push the events through the same map; if the resulting cross-section limits move by more than the quoted map uncertainties (about 25% for efficiencies above 0.5%), the validation's agreement is tune-dependent rather than robust. A second test: generate a model with a deliberately different pseudorapidity distribution and compare map-based limits with a full detector-level recast; if they disagree by more than a factor of a few, the map's folded acceptance assumption fails.","supporting_citations":[{"cited_title":"https : / / www.hepdata.net/record/ins2043503","cited_arxiv_id":null,"evidence_quote":"Provides the efficiency map in structured form and the helper code that defines the bin indexing."}],"review_version":1}