{"id":"d5a906ba-7147-4796-9e97-abbaf4aac1b8","arxiv_id":"2412.19729","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"TooLQit releases FeynRules models for all twelve leptoquark types and a Python calculator, CaLQ, that computes LHC dilepton limits on S1 and U1 couplings for masses from 1 to 5 TeV.","lead":"This paper releases TooLQit, an open-source set of FeynRules model files for all twelve leptoquark types plus a Python calculator named CaLQ that checks whether leptoquark couplings are allowed by LHC dilepton data. It is a practical toolkit for particle physicists studying leptoquark explanations of anomalies and for planning LHC searches.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CaLQ's yes/no output is not benchmarked against any published limit or independent calculator, and with the paper's own 'alpha stage' and 'mimicked' cuts this leaves the central usability claim unverified.","rationale":"TooLQit is a software release, so the appropriate standard is reproducibility and reliability of the delivered tool. The paper provides a clear description of the code and a public repository, which is genuine support, and I did not find an obvious internal contradiction in the FeynRules model listing or in the chi-squared formalism as written. However, the user-facing promise that CaLQ tests whether parameter points are allowed cannot be verified from the manual alone. Every number entering Eq. (9) comes from unseen stored simulation grids and a default 10% systematic that is asserted rather than demonstrated, and the paper explicitly calls CaLQ alpha and says the experimental cuts were only 'mimicked'. The absence of any benchmark comparison against published ATLAS/CMS contours or against HighPT is therefore the single load-bearing gap: it is what would let a user trust that the interpolated cross-sections, binwise efficiencies, and systematic model produce the true limits. The reader's weakest_assumption identifies exactly this gap, and I agree with it. The proposed concrete test—directly comparing CaLQ's one-coupling limits with HighPT and with the published CMS and ATLAS dilepton results—would settle the concern: agreement within the expected accuracy would support the tool and could lift the conditional, while significant disagreement would show that the central usability claim is not yet met. Because the reader's conditional verdict already reflects this uncertainty, I recommend keeping the verdict unchanged.","tokens_in":20449,"tokens_out":11833,"duration_ms":133002,"concrete_test":"Run CaLQ in single-coupling mode for U1 and S1 at M = 1, 2, and 3 TeV for the first-generation couplings, extract the 95% CL bounds, and compare them with the corresponding HighPT predictions and with the CMS 2103.02708 and ATLAS 2002.12223 exclusion contours. If any bound differs by more than about 20% in coupling, or if the allow/exclude status flips for a published benchmark point, the stored simulation grids or the chi-squared systematic model need revision before the tool can be used as advertised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The toolkit's central value is that CaLQ returns correct allow/exclude verdicts for LQ parameter points. That correctness is not established anywhere in the paper: the stored event samples use LO MadGraph5 with MLM matching, Pythia8 showering, Delphes detector simulation, and a default 10% systematic in Eq. (9), while the selection criteria are only described as 'mimicked' from Refs. [56, 57]. No comparison is shown against published ATLAS/CMS exclusion contours or against the independent calculator HighPT. As a result, a user cannot tell whether the 500 GeV interpolation grids and the chi-squared procedure in Eqs. (9)-(11) reproduce the true experimental limits. If the stored efficiencies or the assumed systematic are wrong, every yes/no answer shifts, and the abstract's claim that CaLQ 'can calculate' the LHC limits becomes unsupported. The paper's own alpha-stage label and 'mimicked' wording make this more than a stylistic gap: the reliability of the tool is exactly the load-bearing assertion, and it is left untested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"TooLQit is an open-source toolkit with two components. The first is a set of leading-order FeynRules/UFO model files for all twelve renormalizable scalar and vector leptoquark representations, including their electroweak gauge interactions and a systematic naming convention (Section II). The second is CaLQ, a Python calculator implementing the chi-squared method of Refs. [34, 43, 46] to judge whether a given leptoquark mass and Yukawa-coupling point is allowed by LHC dilepton data (Section III). CaLQ currently supports the singlet scalar S1 and the singlet vector U1 for masses between 1 and 5 TeV, using binned dilepton data from an ATLAS tau-tau search [56] and a CMS ee/mu-mu search [57], with stored LO MadGraph5/Pythia8/Delphes event grids interpolated in 500 GeV steps, a default 10% systematic uncertainty, and an option to ignore resonant contributions. The paper is a manual: it documents model conventions, usage in interactive and non-interactive modes, the chi-squared workflow, and illustrative one- and two-coupling exclusion scans for U1 (Figs. 8-9).","tokens_in":20631,"tokens_out":14231,"duration_ms":136785,"significance":"If the toolkit performs as claimed, it addresses a real community need: it provides, in one open-source framework, FeynRules models for all leptoquark types with the standard electroweak structure (including the photon/Z-gluon-LQ-LQ vertex) and an automated estimator of indirect LHC dilepton limits. The strengths of the paper are the completeness of the model set, the consistent notation, the public repository, the modular design of CaLQ, and the explicit documentation of editable inputs (systematic uncertainty, NLO k-factor, anomalous coupling kappa). The main gap is that the central claim, namely that CaLQ returns correct allow/exclude verdicts, is not established by any comparison with published experimental limits or with an independent calculator. For a tools paper of this type, validation is not an optional extra: users can only trust the output if the stored simulation grids, mimicked cuts, and systematic model have been shown to reproduce known exclusions.","major_comments":[{"comment":"The central usability claim, that CaLQ returns correct allow/exclude verdicts for LQ parameter points, is not validated anywhere in the paper. The stored event samples are generated at LO with MadGraph5, showered with Pythia8, passed through Delphes, and filtered with selection criteria described only as 'mimicked' from Refs. [56, 57]; cross-sections and efficiencies are interpolated from grids spaced by 500 GeV, and a default systematic uncertainty delta = 0.1 enters Eq. (9). No comparison is shown with the published ATLAS/CMS exclusion contours for S1 or U1, with the authors' own published U1 limits in Ref. [43], or with an independent calculator such as HighPT [58, 59]. Because the purpose of the tool is precisely to decide whether a parameter point is excluded, the absence of a closure test leaves the abstract's claim that CaLQ 'can calculate the LHC limits' unsupported. I recommend adding a validation section that reproduces at least the U1 limits of Ref. [43] and the experimental contours for a few benchmark masses, and ideally compares CaLQ output with HighPT for the overlapping masses.","section":"Section III, Eqs. (9)-(11), Figs. 8-9"},{"comment":"The origin of the SM background N_b_SM in Eq. (10) is not stated. The chi-squared in Eq. (9) compares N_b_Theory = N_b_LQ + N_b_SM with the observed N_b_Data, and every verdict depends on what N_b_SM contains (e.g., the HepData background columns versus an independent simulation) and on whether its uncertainty is included in Delta-N. The manuscript should specify the source of N_b_SM, describe how it is retrieved and stored, and state whether the 10% systematic is intended to cover its normalization uncertainty; as written, the definition is incomplete and the results cannot be reproduced without inspecting the code in the repository.","section":"Section III, Eq. (10)"},{"comment":"The default systematic model, Delta-N_b_syst = delta * N_b_Data with delta = 0.1 treated as uncorrelated across bins, is a free input whose impact on the limits is not quantified. The yes/no verdicts shown in Figs. 8-9 will shift with delta, and the paper offers neither a delta-dependence study nor external evidence that 0.1 is the right value for these particular analyses. I ask the authors to show the effect of varying delta on at least one representative limit curve, or to absorb this sensitivity check into the validation section requested above. Without such a check, the robustness of the calculator's exclusion boundaries is unknown.","section":"Section III, Eq. (9) and Section III.B"}],"minor_comments":[{"comment":"Table I omits the barred scalar and vector states S-bar_1 and U-bar_1, even though they appear in Eq. (1) and Table II; the authors should either add their Yukawa interaction terms or explicitly state that the table excludes them.","section":"Table I"},{"comment":"The Monte Carlo code assigned to the scalar S-bar_1 is '4210212', whose third digit (1) denotes a vector in the scheme stated in Section II.A; to follow the stated convention consistently, the code should be '4200212'.","section":"Table II and Section II.A"},{"comment":"In the naming-convention bullet, the text 'r (for R2 and eR1)' should read 'eR2'.","section":"Section II.A"},{"comment":"In the sentence listing the six scalar LQs, the second 'S1' (the charge -2/3 state) loses its bar in the typesetting; the bar should be visible so that the list agrees with Eq. (1).","section":"Section II"},{"comment":"The directory path 'TooLQit/CaLQ/Version_X.Y.Z' is a placeholder; the manuscript should cite the concrete version of the code (or the repository state) used to produce Figs. 8-9.","section":"Section III.A"},{"comment":"At least one panel of Fig. 8 should be overlaid with the corresponding published exclusion limit or the limit from Ref. [43] so that the reader can judge the scale of CaLQ's grey regions; without any reference curve, the displayed exclusions cannot be assessed.","section":"Figure 8"},{"comment":"The sentence referring to 'the scipy.optimize() function' should name the actual routine used (e.g., scipy.optimize.minimize) and state whether the minimization is constrained to the [-3.5, 3.5] interval used for the coupling evaluation.","section":"Section III.C"}],"recommendation":"major_revision","confidential_remarks":"The paper is essentially a code manual, so the verification burden rests on demonstrating that the code reproduces known limits. A self-validation is trivially available because the authors' own Refs. [34, 43, 46] already published U1 and S1 limits with the same chi-squared method; the absence of such a comparison is conspicuous and should be addressed before acceptance. I would also note that the paper combines data from two different experiments (ATLAS tau-tau and CMS ee/mu-mu) with 'mimicked' cuts; even after validation at the level of final limits, the documentation should warn users about the resulting caveats. The manuscript fits the journal's scope as a computational tools contribution provided the validation gap is closed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: the new content here is the software, not the physics. The FeynRules model files for all twelve LQs—with electroweak gauge vertices and a coherent naming scheme—plus the public CaLQ code, are a genuinely useful resource for leptoquark phenomenology. The illustrative U1 limits repeat the authors' own earlier analyses, so don't expect new physics results.\n\nWhat is good: the paper is honest about what it is. It states that CaLQ is at alpha stage, describes the chi-squared workflow in enough detail to reproduce it, and ships the code publicly. Including the gamma/Z-g-LQ-LQ vertices in the LO models is a real plus. As a manual, the paper works.\n\nThe soft spot is exactly where the stress-test note points: CaLQ's yes/no answers are not validated. The selection cuts are 'mimicked' from ATLAS/CMS, the simulation is LO with Delphes and 500 GeV interpolation grids, and the default 10% systematic is assumed. No comparison is shown with published ATLAS/CMS exclusion contours or with HighPT. Since the calculator's whole job is to reproduce those limits, the lack of a single benchmark plot leaves the central usability claim unverified. This is not a fatal flaw in the math—the chi-squared method is standard and the code can be checked—but it is a load-bearing evidential gap, and the paper's own alpha-stage label makes it harder to wave off. A user genuinely cannot tell whether the stored efficiencies are accurate.\n\nOn self-citation: the method is from their earlier papers. I don't count that against them; the method is well documented, and the toolkit's contribution is the automation, not the technique. A minor complaint: the paper names the repository but no commit hash or version stamp, so reproducibility of the exact code is harder to pin down.\n\nBottom line: this paper deserves a serious referee. It is a useful tool that probably works, but the authors should be asked to add a validation section comparing CaLQ output with published limits or an independent calculator before the tool is used as a black box. For a pheno group working on LQs, this is worth a reading group slot. I would cite the FeynRules models; I would hold off citing CaLQ limits until the benchmark exists.","headline":"Useful, honest LQ software paper whose central claim—CaLQ reproduces LHC limits—is not yet validated; referees should ask for a benchmark.","tokens_in":21234,"tokens_out":3090,"would_cite":true,"duration_ms":32583,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper presents TooLQit, an open-source toolkit with model files for all twelve leptoquark types and a calculator that tests S1 and U1 points against LHC dilepton limits.","keywords":["leptoquarks","LHC phenomenology","dilepton tails","chi-squared limit estimation","S1 scalar leptoquark","U1 vector leptoquark","beyond the Standard Model","Monte Carlo event generation"],"falsifier":"Take a set of coupling-mass points that the LHC experiments have already excluded at 95% confidence in their public dilepton contours, run each through CaLQ, and count how many return 'yes'. If a sizable fraction are flagged as allowed, or if regenerating the stored event samples at an off-grid mass shifts binned predictions beyond the quoted uncertainties, the calculator's claim to estimate the indirect limits is falsified.","tokens_in":1851,"feed_emoji":"🧰","tokens_out":4904,"duration_ms":152160,"temperature":0.7,"pith_summary":"TooLQit is an open-source toolkit for studying leptoquarks—hypothetical particles that couple quarks to leptons—at the LHC. The paper's central claim is that the toolkit supplies computer-ready leading-order model files for all twelve renormalizable leptoquark types, including their electroweak gauge interactions, and a calculator, CaLQ, that tells a user whether a chosen mass and set of leptoquark–quark–lepton couplings is still allowed by the high-mass dilepton data. CaLQ works on a chi-squared method: it compares measured event counts in electron, muon, and tau-pair channels with predictions built from stored cross-sections and cut efficiencies, minimises over the active couplings, and returns yes or no at the requested significance. The paper offers the toolkit as a proof-of-principle of a unified, modular platform for leptoquark phenomenology. If it works as described, model-builders can screen large parameter spaces and simulate leptoquark signals without rebuilding models or rederiving constraints from scratch.","feed_headline":"Twelve leptoquark models, one toolkit, built-in LHC limit checks","feed_subtitle":"Model files for all twelve leptoquark types, plus a calculator that tests S1 and U1 against dilepton data.","key_machinery":"The load-bearing mechanism is the chi-squared limit calculator built on a factorised parametrisation of BSM event counts. For each dilepton channel and bin, CaLQ writes the predicted number of events as the Standard Model background plus three leptoquark contributions—pair production, single production, and t-channel exchange interfering with the Drell-Yan background—and expresses the coupling dependence of the non-resonant piece through terms proportional to $\\lambda_i^2$ and $\\lambda_i^2 \\lambda_j^2$ with precomputed cross-sections and efficiencies. It then minimises the total chi-squared in the space of the chosen couplings (taken real, varied between $-3.5$ and $3.5$, with multiple starting points) and decides allowed or excluded by comparing the input point's $\\Delta\\chi^2$ with the $1\\sigma$ or $2\\sigma$ threshold. A second piece of machinery is the model-file infrastructure itself: a systematic naming and coupling convention, plus auxiliary fields, that lets electroweak gauge vertices be grafted onto every leptoquark representation for event generation.","core_discovery":"The paper claims that all twelve renormalizable leptoquark models—six scalars and six vectors—can be packaged into standard computer-readable model files that include the electroweak gauge interactions, not just the QCD ones, and that the two singlet cases S1 and U1 can be fed into an automated calculator which turns the LHC's published high-mass dilepton distributions into exclusion statements. That calculator, CaLQ, is an automation of the chi-squared limit-setting technique from the authors' earlier work: for a fixed mass it reads stored cross-sections and bin-by-bin efficiencies for pair, single, and indirect production, interpolates them between mass points, forms the predicted event count in each dilepton bin, minimises the chi-squared over the active couplings, and accepts or rejects the input point by the resulting $\\Delta\\chi^2$. The paper demonstrates the workflow with one- and two-coupling scans for the U1 vector leptoquark and describes the release as a proof-of-principle of a unified, modular leptoquark framework.","pith_inferences":["The default 10% systematic uncertainty and the 500 GeV interpolation grid are likely to dominate the error budget; a user who tests a point at neighbouring masses or with the systematic error set to 5% and 15% will get a sense of how robust the yes/no boundary actually is.","The same machinery transfers naturally to any BSM scenario that alters dilepton tails, such as effective-field-theory Wilson coefficients; the paper notes this possibility but does not implement it.","Because the parametrisation assumes real couplings, points with complex phases may be misclassified; the paper argues the LHC data are largely insensitive to this, so the practical effect is probably small.","The planned inclusion of NLO-QCD models and direct-search limits would change the reach substantially; until then, leading-order cross-sections with a flat k-factor for scalar pair production are the main approximation."],"forward_implications":["A model-builder can test an S1 or U1 parameter point against electron, muon, and tau-pair tails in minutes, for any mass from 1 to 5 TeV, instead of rerunning the full simulation chain.","The model files make it straightforward to generate leading-order signals for any of the twelve leptoquarks, including processes with photon/Z-gluon-LQ-LQ vertices that are often omitted but can matter at high electric charge.","Because CaLQ is modular and the chi-squared procedure is generic, the same stored-data approach can be extended to other leptoquark representations, mixed-flavour final states, and lepton-plus-missing-energy data.","The non-interactive mode allows whole grids of parameter points to be screened, so points allowed by flavour or dark-matter constraints can be checked against LHC dilepton bounds in one pass."],"supporting_citations":[{"why":"Introduces the interference-sensitive chi-squared exclusion method for S1 that CaLQ automates.","marker":"[34]"},{"why":"Generalises the method to the U1 vector leptoquark and supplies the exact limit-setting technique CaLQ implements.","marker":"[43]"},{"why":"Supplies the tau-pair transverse-mass binned data used in the chi-squared sum.","marker":"[56]"},{"why":"Supplies the electron-pair and muon-pair invariant-mass binned data used in the chi-squared sum.","marker":"[57]"},{"why":"Catalogue of the twelve renormalizable leptoquark representations that the model-file set covers.","marker":"[14]"},{"why":"Provides the model-generation software used to produce the model files.","marker":"[55]"},{"why":"Provides the output format that carries the models into matrix-element event generators.","marker":"[60]"},{"why":"Generates the leading-order event samples whose cross-sections and efficiencies populate CaLQ's stored grids.","marker":"[62]"}],"fun_headline_variants":["All 12 leptoquarks in one toolkit with LHC limits","Leptoquark models unified: CaLQ tests S1, U1 vs LHC","One toolkit for all leptoquarks: auto LHC limit checks","TooLQit: every leptoquark model plus automated LHC limits","Automated leptoquark limits: all 12 models, CaLQ included"],"cache_read_input_tokens":23424,"weakest_assumption_plain":"The calculator's yes/no answers rest on the stored leading-order event samples, the 500 GeV interpolation grid, a default 10% systematic uncertainty, and the assumption that these reproduce the true experimental limits; the paper does not compare CaLQ output with published exclusion contours or with an independent calculator, so that fidelity is unverified.","fun_headline_variants_meta":{"raw":{"variants":["All 12 leptoquarks in one toolkit with LHC limits","Leptoquark models unified: CaLQ tests S1, U1 vs LHC","One toolkit for all leptoquarks: auto LHC limit checks","TooLQit: every leptoquark model plus automated LHC limits","Automated leptoquark limits: all 12 models, CaLQ included"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000261,"raw_usage":{"total_tokens":1596,"prompt_tokens":954,"completion_tokens":642,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":540}},"tokens_in":570,"tokens_out":642,"duration_ms":31108,"temperature":1.0,"reasoning_tokens":540,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:55:20.224761+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of coupling-mass points that the LHC experiments have already excluded at 95% confidence in their public dilepton contours, run each through CaLQ, and count how many return 'yes'. If a sizable fraction are flagged as allowed, or if regenerating the stored event samples at an off-grid mass shifts binned predictions beyond the quoted uncertainties, the calculator's claim to estimate the indirect limits is falsified.","supporting_citations":[],"review_version":1}