{"id":"69bc361e-e87e-4f68-8d18-b4a5158f5102","arxiv_id":"2505.05532","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"WOMBAT, a quantized ML trigger for boosted H to b bbar jets, reaches a 1 kHz CMS trigger rate at jet pT around 140 to 147 GeV, about 40 to 47 GeV lower than the Single Jet 180 baseline.","lead":"WOMBAT is an ML-based Level-1 trigger for CMS that identifies boosted Higgs to bottom quark jets from calorimeter data, reaching the 1 kHz rate budget at lower jet momentum than the standard trigger. A quantized version fits on a Virtex-7 FPGA in synthesis, but at 22 clock cycles it still misses the stated 14-cycle latency budget.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"W-AM latency of 22 clock cycles violates the paper's own 14-cycle L1T budget, so the central deployability claim is internally contradicted.","rationale":"The reader's weakest_assumption focused on the normalization factor Nh0/Nh1 in Eq. 73, which is a valid concern about the rate comparison. However, the single most load-bearing issue is the latency contradiction: the paper explicitly defines the L1T processing budget as 14 CC, then reports W-AM at 22 CC and admits none of the algorithms meet the requirement. This directly invalidates the headline claim that W-AM is deployable within L1T constraints. The rate normalization affects the magnitude of the threshold improvement, but the latency issue undermines the hardware feasibility claim entirely, independent of any physics comparison. The paper's honest acknowledgment of the failure is a credit, but the abstract and concluding claims overstate the result. A CONDITIONAL verdict remains appropriate because the work retains value as a proof-of-concept for Phase-2 and as a detailed architecture study, provided the deployability claim is corrected or the latency budget is explicitly revisited. The reader's rationale did mention the latency tension, so agreement is partial rather than full.","tokens_in":51084,"tokens_out":2073,"duration_ms":24702,"concrete_test":"Check Table 7 and Chapter V Section 6: confirm that W-AM's reported latency is 22 CC and that the stated L1T budget is 14 CC. Re-run the HLS synthesis with the same configuration and inspect the latency report; if the reported latency is 22 CC, then W-AM misses the 87.5 ns budget by 8 CC. Optionally, check whether the 14 CC budget is the correct Phase-1 L1T calorimeter trigger latency by consulting the CMS L1T technical design documentation; if the true budget is larger, the deployability claim would need to be re-evaluated against that number, but the paper's own stated constraint is already violated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim that W-AM is FPGA-deployable is contradicted by the paper's own numbers. Chapter V Section 6 states that the L1T processing budget is 14 clock cycles (87.5 ns at 6.25 ns/CC), and immediately concedes: \"As shown in Table 7, none of the algorithms meet the L1T FPGA processing time requirement of < 87.5 ns.\" Table 7 reports W-AM at 22 CC (137.5 ns) for the DATAFLOW implementation, 8 CC over budget, with JEDI even worse at 56 CC. The abstract nonetheless claims W-AM \"meets resource constraints with a pre-place-and-route latency of 22 clock cycles (137.5 ns)\", which is misleading because 22 CC does not meet the stated latency constraint. This is not a matter of disagreeing with the CMS latency definition; the paper itself defines the 14 CC budget. The pre-place-and-route caveat does not rescue the claim: post-placement-and-route latency in clock cycles does not decrease, and the synthesis report already gives cycle-true latency. Therefore the central claim that W-AM is deployable in the Phase-1 L1T is false by the paper's own metric, even if the physics performance proof-of-concept remains of interest.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This thesis-style manuscript presents WOMBAT, a machine-learning-based Level-1 trigger algorithm for boosted H→bbbar jet identification at CMS, together with its FPGA implementation. WOMBAT comprises a high-performance teacher model (W-MM) and a quantized, FPGA-synthesizable apprentice model (W-AM), and is benchmarked against the existing Single Jet 180 trigger and a hand-crafted rule-based algorithm called JEDI. Rates are evaluated with 2023 CMS ZeroBias data, efficiencies with a dedicated H→bbbar Monte Carlo sample, and firmware cost with HLS synthesis on a Xilinx Virtex-7 device. The headline claims are that W-MM and W-AM sustain a 1 kHz trigger rate at offline jet pT thresholds of 146.8 GeV and 140.4 GeV, respectively, compared with 187.4 GeV for Single Jet 180, and that W-AM fits on the target FPGA with a 22 clock-cycle pre-place-and-route latency, while JEDI needs 56 clock cycles and excessive resources. The manuscript explicitly frames WOMBAT as a proof-of-concept for Phase-2 L1 triggers.","tokens_in":51331,"tokens_out":7149,"duration_ms":79696,"significance":"If the quantitative claims survive scrutiny, the paper would be a useful contribution to the L1 machine-learning trigger literature: it pairs a concrete physics target with real ZeroBias rate evaluation, a complete quantized model description, firmware latency/resource numbers, and links to public repositories, which is a strong level of reproducibility for this type of work. The comparison with a deterministic baseline (JEDI) is also valuable because it demonstrates that simple rule-based designs do not automatically dominate learned triggers. However, the central deployability statement is contradicted by the paper's own latency budget, and the rate comparison depends on an unexplained normalization factor. Those two issues must be resolved before the headline threshold improvements can be taken as a reliable physics result.","major_comments":[{"comment":"The paper defines the L1T FPGA processing budget as 14 clock cycles (87.5 ns at 6.25 ns/CC) and then states explicitly that none of the algorithms meet that requirement. Table 7 reports W-AM at 22 CC (137.5 ns), i.e., 8 CC over budget, and JEDI at 56 CC. The abstract nonetheless presents the 22 CC latency as evidence that W-AM meets resource constraints, and the concluding summaries describe W-AM as deployable. By the paper's own metric, W-AM does not meet the Phase-1 L1T latency budget. This is an internal contradiction in the central deployability claim, not a matter of external disagreement. Please revise the deployment wording so that W-AM is described as exceeding the Phase-1 latency budget while remaining a hardware-feasibility proof-of-concept, or provide a concrete, reasoned argument for why the 14 CC budget does not apply to this algorithm in its intended position in the L1T chain.","section":"Ch. V, Sec. 6, Table 7; also abstract"},{"comment":"Equation (73) multiplies the cumulative event fraction N_{>=pT(i)}/N_total by the factor N_{h0}/N_{h1}, described as a correction for differences in normalization between the WOMBAT and Single Jet 180 rate histograms. If both algorithms process the same ZeroBias sample, the raw event fractions should already be directly comparable; if the n-tuples have different total counts because of different skim conditions or processing paths, the ratio can compensate for those differences but can also imprint them directly on the derived pT thresholds. The manuscript provides no justification that the ratio is unbiased, no test of its stability, and no uncertainty attached to it. Since the headline 40.6 GeV and 47.0 GeV threshold improvements are computed from this formula, this normalization factor is load-bearing. Please either derive the factor from the sample definitions, demonstrate that the conclusion is unchanged on a common subsample, or remove the factor and recompute the rates.","section":"Ch. IV, Sec. 7, Eq. (73)"},{"comment":"The abstract states that W-MM achieves its lower pT threshold while maintaining comparable signal efficiency to Single Jet 180. Under the standard DeltaR < 0.4 matching used in the analysis, Table 3 reports epsilon(pT) = 0.71 for W-MM versus 0.91 for Single Jet 180 at the 1 kHz threshold. That is a substantial efficiency loss, not a comparable one. Under DeltaR < 0.8 the numbers are closer (0.89 vs 0.95), but the abstract does not state this condition. The text should either qualify the efficiency comparison by the matching criterion or soften the claim. The same precision is needed in the conclusion where W-AM is described as deployable; as noted in the first major comment, that wording is contradicted by the latency table.","section":"Abstract and Ch. V, Sec. 3, Table 3"},{"comment":"The resource utilization figures are HLS synthesis estimates only, before placement and routing. The manuscript asserts that HLS estimates are usually pessimistic for 7-series devices and therefore adequate for algorithmic comparison, but the central claim that W-AM 'meets resource constraints' on a specific FPGA is stronger than what a pre-P&R report can establish. Please explicitly label the resource numbers as pre-implementation estimates, and either provide post-implementation results or explain in the deployment discussion why the estimates are sufficient for the intended conclusion. This is a correctness-of-wording issue rather than a fundamental flaw, but it feeds the same over-claim identified in the first major comment.","section":"Ch. V, Sec. 6, Table 8"}],"minor_comments":[{"comment":"The two equations defining the masks r_eta and r_phi are both labeled r_eta; the second line should define r_phi.","section":"Ch. IV, Sec. 4, after Eq. (70)"},{"comment":"The conversion factor is written as 40x10^6 / 10^3 [kHz]; the dimensional logic would be clearer if the text stated that it converts a per-bunch-crossing fraction to kHz through the 40 MHz bunch-crossing rate.","section":"Ch. IV, Sec. 7, Eq. (73)"},{"comment":"The notation in Eq. (27) uses a comma inside the max function arguments in a way that is easy to misread; a clearer formulation would separate the two pooled entries explicitly.","section":"Ch. IV, Sec. 2.2.1, Eq. (27)"},{"comment":"For reproducibility, the repository list would benefit from commit hashes or version tags; the current table identifies projects but not the exact code versions used for the numbers in Tables 7 and 8.","section":"Appendix F, Table 12"}],"recommendation":"major_revision","confidential_remarks":"The paper has solid empirical content and an unusually complete FPGA/rate/efficiency narrative for a thesis-based submission. The main concern is that the abstract and conclusions over-claim deployability in a way that the paper's own Table 7 explicitly contradicts. That is fixable by reframing the contribution as a hardware-feasibility proof-of-concept and by tightening the efficiency and rate-normalization statements. I do not see a need to reject, provided the authors address the latency contradiction and the Eq. (73) normalization factor in a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read of the WOMBAT thesis. The real content: a complete ML-based L1 calorimeter trigger chain for boosted H→bb, from architecture through knowledge distillation to a QKeras/HLS FPGA implementation, evaluated on 2023 ZeroBias data and MC. The specific combination—embedded deterministic autoencoder with circular φ padding, teacher-student quantization, and a working firmware prototype on the Virtex-7—is genuinely new as an integrated system. The headline numbers (1 kHz rate at 146.8 GeV for W-MM, 140.4 GeV for W-AM, versus 187.4 GeV for Single Jet 180) are internally consistent with the rate plots, and the detailed firmware optimization discussion (DATAFLOW approach, buffer compaction, reuse factors) is a real engineering contribution. Credit where due: this is reproducible work with repositories listed and enough detail to reimplement.\n\nThe soft spots are real but not fatal. The most important: the paper claims W-AM is deployable while Chapter V Section 6 states none of the algorithms meet the <87.5 ns budget; Table 7 shows W-AM at 22 CC (137.5 ns). The abstract calls out 22 cycles as if it satisfies the constraint. That is a contradiction by the paper's own metric, and the stress-test note is right. The pre-place-and-route caveat does not rescue the claim; cycle-true latency does not shrink after P&R. Second, the rate formula (Eq. 73) includes a normalization factor Nh0/Nh1 with no justification; if the n-tuple acceptances differ between WOMBAT and Single Jet 180, the 40 GeV threshold improvement could be partly an artifact. Third, W-AM efficiency at ΔR<0.4 is 0.53 at threshold, and the paper honestly shows the efficiency drop at high pT due to fixed two-jet output. That limits the physics claim but does not invalidate the extended-coverage argument. Missing: benchmarks against earlier ML-based L1 taggers beyond the rule-based baseline.\n\nWho this is for: people working on ML trigger firmware or boosted H→bb triggering. It is a solid proof-of-concept, not a deployed trigger. The deployability wording needs fixing and the rate normalization needs derivation; both are referee-level revisions. I would send it to review, not desk reject.","headline":"Solid ML trigger proof-of-concept with a real firmware implementation, but the deployability claim is internally contradicted by the 22-cycle latency and the rate normalization needs justification.","tokens_in":51900,"tokens_out":2004,"would_cite":true,"duration_ms":21859,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An FPGA neural network trigger can select boosted H→bbar jets at a 1 kHz rate with a jet pT threshold about 40 GeV lower than the standard CMS Single Jet 180 trigger.","keywords":["Level-1 trigger","boosted H to b bbar","jet tagging","FPGA","deep neural network","quantized CNN","knowledge distillation","CMS calorimeter trigger"],"falsifier":"Recompute the 1 kHz rate crossing using the raw event counts without the normalization factor, or split the ZeroBias sample into independent run periods and require the threshold difference to reproduce in each; a movement larger than the ±5.5 GeV bin width would falsify the claim as stated.","tokens_in":50836,"feed_emoji":"⚛️","tokens_out":6882,"duration_ms":68495,"temperature":0.7,"pith_summary":"This paper tries to establish that a machine-learned trigger running on the FPGAs of the CMS Level-1 calorimeter trigger can identify Lorentz-boosted H→bbar jets using only trigger-primitive energy grids. The central result is that the WOMBAT Master Model reaches a 1 kHz trigger rate at an offline jet pT of 146.8 GeV, while the standard Single Jet 180 trigger needs 187.4 GeV for the same rate, and the FPGA-synthesizable Apprentice Model reaches 140.4 GeV. If correct, this would let CMS collect boosted Higgs and di-Higgs events in a kinematic window roughly 40 GeV wider than the present baseline, without spending additional trigger bandwidth. The paper also shows the apprentice model fits on a Virtex-7-class FPGA with 22 clock cycles of pre-place-and-route latency, establishing a proof of concept for machine-learned jet tagging at the first trigger level.","feed_headline":"Neural trigger catches boosted Higgs jets 40 GeV below baseline","feed_subtitle":"At the same 1 kHz rate, the FPGA trigger passes jets above 146.8 GeV where CMS's standard trigger needs 187.4 GeV.","key_machinery":"The carrying mechanism is a teacher–student knowledge-distillation pair sharing one input representation: the W-MM uses an Embedded Deterministic Autoencoder (EDA) that treats the azimuthal angle φ as cyclic—circularly padding that dimension while keeping η at full resolution—so the network learns jet substructure that wraps across the detector boundary. The distilled W-AM is an 8-bit quantized CNN with a custom pT-threshold layer that subtracts 30 GeV from each trigger region before convolution, filtering low-energy noise at no latency cost; its fixed two-jet output is what keeps it small enough for the FPGA. JEDI, the comparison baseline, is a fully rule-based pipeline of pileup-subtracted 3×3 energy sums, shape masks, and a 64-element bitonic sorter.","core_discovery":"The claim, stated on the paper's own terms, is that WOMBAT—a convolutional-neural-network trigger built around an Embedded Deterministic Autoencoder and a distilled, 8-bit-quantized apprentice—can localize boosted H→bbar jets directly from a 14×18 grid of calorimeter trigger regions. W-MM achieves the 1 kHz rate at 146.8 GeV and W-AM at 140.4 GeV, compared with 187.4 GeV for Single Jet 180, with W-MM maintaining signal efficiency comparable to the baseline and W-AM accepting a lower-efficiency trade for hardware feasibility. The FPGA implementation of W-AM is reported at 22 clock cycles (137.5 ns) pre-place-and-route with resource usage within the target device, while the rule-based JEDI baseline needs 56 cycles and excessive resources. The paper positions the result as a Phase-1 proof of concept for Phase-2 triggers, where larger latency budgets and finer granularity would allow the same method to be deployed online.","pith_inferences":["Editorial inference: The 40 GeV threshold gain should be read as conditional on the normalization correction in the rate formula; if the raw event counts for the two triggers differ because of acceptance rather than trigger path, the gain could shrink or vanish.","Editorial inference: The EDA's circular-φ design is a reusable ingredient for any online trigger that must find objects crossing the detector boundary; one testable extension is to apply the same teacher–student recipe to boosted W/Z or top jets, where the substructure signatures differ but the boundary problem is identical.","Editorial inference: A stricter test of the claimed gain would be to run W-AM and Single Jet 180 on the same events with identical n-tuple acceptances, or to measure the rate on a second independent sample; the paper does not report such a cross-check, so the numerical threshold difference is not yet settled."],"forward_implications":["At a fixed 1 kHz per-trigger rate, the ML trigger accepts events with offline jet pT as low as about 147 GeV, opening the 147–187 GeV window that the standard single-jet trigger cannot record.","Because W-AM's efficiency is capped by its two-jet output in events with three or more jets, a logical OR of W-AM at low pT with Single Jet 180 at high pT would recover efficiency across the full range without exceeding the rate budget.","The 22-cycle pre-place-and-route latency exceeds the Phase-1 14-cycle budget, so the immediate claim is not deployment but feasibility: quantization, thresholding, and pipeline restructuring bring an ML tagger within reach of current FPGA capacity.","JEDI's 56-cycle latency and heavy resource use indicate that rule-based boosted-Higgs tagging is the harder path to fit in the Level-1 hardware, strengthening the case for learned features."],"supporting_citations":[{"why":"Supplies the NLO matrix-element signal sample used for training and efficiency evaluation.","marker":"[63]"},{"why":"Provides parton showering and hadronization for the same Monte Carlo sample.","marker":"[65]"},{"why":"Emulates the detector response that produces the calorimeter trigger primitives.","marker":"[66]"},{"why":"Supplies the 8-bit quantization layers used to build the apprentice model.","marker":"[71]"},{"why":"Converts the trained quantized model into high-level synthesis code for the FPGA implementation.","marker":"[72]"},{"why":"Defines the target Virtex-7 FPGA whose resources and clock constrain the design.","marker":"[75]"},{"why":"Justifies comparing pre-place-and-route latency and resource estimates across algorithms.","marker":"[79]"}],"fun_headline_variants":["Neural L1 trigger: 40 GeV lower threshold for boosted Higgs","FPGA neural trigger tags boosted Higgs at 1 kHz rate","WOMBAT: 22-cycle neural trigger, 40 GeV threshold drop","CMS L1 neural net: 40 GeV gain, fits FPGA budget"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claimed threshold gain rests on a normalization factor that rescales the WOMBAT and Single Jet 180 rate histograms to a common event count; if that factor is correcting for different event selection rather than plain luminosity, the 40 GeV improvement could be an artifact.","fun_headline_variants_meta":{"raw":{"variants":["Neural L1 trigger: 40 GeV lower threshold for boosted Higgs","FPGA neural trigger tags boosted Higgs at 1 kHz rate","WOMBAT: 22-cycle neural trigger, 40 GeV threshold drop","CMS L1 neural net: 40 GeV gain, fits FPGA budget"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000805,"raw_usage":{"total_tokens":3648,"prompt_tokens":1169,"completion_tokens":2479,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":785,"completion_tokens_details":{"reasoning_tokens":2400}},"tokens_in":785,"tokens_out":2479,"duration_ms":19357,"temperature":1.0,"reasoning_tokens":2400,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:06:53.956567+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the 1 kHz rate crossing using the raw event counts without the normalization factor, or split the ZeroBias sample into independent run periods and require the threshold difference to reproduce in each; a movement larger than the ±5.5 GeV bin width would falsify the claim as stated.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Emulates the detector response that produces the calorimeter trigger primitives."},{"cited_title":"(n.d.-a)","cited_arxiv_id":null,"evidence_quote":"Supplies the 8-bit quantization layers used to build the apprentice model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Converts the trained quantized model into high-level synthesis code for the FPGA implementation."},{"cited_title":"Xilinx Sales | FPGAs (Field Programmable Gate Array) | CPLDs (Complex Programmable Logic Devices)","cited_arxiv_id":null,"evidence_quote":"Defines the target Virtex-7 FPGA whose resources and clock constrain the design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies comparing pre-place-and-route latency and resource estimates across algorithms."}],"review_version":1}