{"id":"822dd198-ad11-4f5c-897f-54cd0496bc18","arxiv_id":"2501.04845","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Simulations show a graph-neural-network trigger on FPGAs can identify beauty-quark decays at sPHENIX with 97% accuracy, and early firmware runs at 505 ns to 9.2 microseconds for simplified models.","lead":"This project report describes putting AI models on FPGAs so the sPHENIX detector at RHIC can find rare heavy-quark decays in real time and keep data it currently throws away. The same approach could later be used at the future Electron-Ion Collider.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported sub-component latencies do not establish the 10 us end-to-end budget: the high-accuracy track-based BGN-ST pipeline has not been run as a whole on the FELIX-712 board, and the one measured stage already consumes 8.82 us on a larger FPGA.","rationale":"The paper is best read as a status report, and its sub-component measurements are honest and useful: an 8.82 us edge-classification latency, a 9.2 us FlowGNN implementation of a simplified hit-based model, and a 505 ns hls4ml version are concrete progress. The problem is that none of these numbers, individually or together, demonstrates the central real-time claim. The advertised physics gain comes from the attention-based BGN-ST model (97.38% accuracy on beauty decays), but the only implementation results for BGN-ST are for a single stage, the edge-candidate classifier, and that stage was measured on an Alveo U280 rather than the target FELIX-712. Since 8.82 us is already 88% of the 10 us stretched target, the remaining stages—event building, edge candidate generation, track construction, momentum estimation, and the BGN-ST vertex/trigger detection itself—would have to fit into a small residual budget, all on a board with roughly half the resources. The 9.2 us and 505 ns FPGA results are explicitly for the hit-based GarNet model, which the paper reports as less accurate (90.57% vs 97.38%); they cannot be used to certify the track-based pipeline. The paper's own Section 5 limits the current integration claim to 'the simplified hit-based model' and places full-system testing later. If the full BGN-ST pipeline is later shown to meet the latency budget on FELIX-712, the central claim would be substantially supported; until then, the appropriate assessment is conditional on that demonstration. I therefore agree with the reader's weakest-assumption identification and see no reason to change the CONDITIONAL verdict.","tokens_in":5172,"tokens_out":4276,"duration_ms":40279,"concrete_test":"Synthesize the complete track-based BGN-ST pipeline (hit decoding/clustering, event building, edge candidate generation, edge classification, track construction, pT prediction, trigger detection) on the FELIX-712 board using realistic sPHENIX event data, and measure sustained latency at 3 MHz input. If the measured end-to-end latency exceeds ~10 us, or if the design does not fit within FELIX-712 resources, the 90% luminosity-recovery claim for the attention-based pipeline fails. An acceptable intermediate check is a per-stage latency/resource breakdown on FELIX-712 with the same clock frequencies and realistic input sizes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 1 defines the project's stretched goal as a trigger within ~10 us, because the TPC buffers hold ~30 us. The full track-based pipeline (Section 2) includes hit decoding/clustering, event building, edge candidate generation, edge classification, track construction, transverse-momentum prediction, and trigger detection. Section 4 reports that the second version of the edge-candidate classifier achieves 8.82 us latency at 285 MHz on an Alveo U280, which the paper calls 'approximately twice as big as FELIX-712'. This single stage therefore leaves at most ~1.2 us of the 10 us target for all remaining stages, before accounting for the resource reduction on the smaller FELIX-712 board. The other end-to-end FPGA numbers reported in Section 4—9.2 us FlowGNN and 505 ns hls4ml—are for the simplified hit-based GarNet model, not for the track-based BGN-ST model that produces the 97.38% accuracy in Section 3.2. Section 5 confirms that only the simplified hit-based firmware pieces are being combined, and that full-system testing is expected by the end of the year. Consequently, the central claim that an FPGA-based BGN-ST trigger can save a large fraction of the lost 90% luminosity is not yet supported by an end-to-end implementation; it rests on extrapolating a sub-component latency from a larger device and on unstated assumptions about pipelining the remaining stages.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper, a proceedings contribution from ICHEP 2024, describes an ongoing R&D project to build an FPGA-based trigger for sPHENIX that uses streaming tracking data to select rare beauty-decay events within the TPC's ~30 microsecond buffer. The proposed pipeline reconstructs tracks from silicon hits via graph-neural-network edge classification, estimates transverse momentum from track curvature, and uses a Bipartite Graph Network with Set Transformers (BGN-ST) for trigger decisions. The authors report 97.38% accuracy for beauty detection on simulated data and FPGA latencies for several sub-components using FlowGNN and hls4ml, including 8.82 us for one edge-classification stage on an Alveo U280 and 505 ns for a simplified hit-based model.","tokens_in":5354,"tokens_out":3976,"duration_ms":35197,"significance":"If the complete pipeline meets its 10 us end-to-end latency on the FELIX-712 board while preserving the simulated accuracy, it would enable sPHENIX to capture beauty-enriched events from the 90% of luminosity currently not saved, a major gain for the heavy-flavor program. The work combines a state-of-the-art attention-based GNN with hardware implementation and explicitly compares against baseline models. The manuscript's honesty about its current status (Section 5: only simplified hit-based firmware is being combined; full-system testing expected by end of year) is a strength, as is the direct measurement of sub-component latencies. However, the headline claim of \"huge improvement\" is forward-looking and rests on extrapolation.","major_comments":[{"comment":"The end-to-end 10 us trigger budget is not yet demonstrated. The only track-based FPGA result is the edge-candidate classification at 8.82 us on an Alveo U280, which the paper itself states is approximately twice as large as the target FELIX-712; this leaves under 1.2 us for all remaining stages (event building, remaining track construction, momentum prediction, trigger detection) before accounting for resource scaling. The 9.2 us and 505 ns results are for the simplified hit-based GarNet model, not for the BGN-ST model that produces the 97.38% accuracy. Section 5 confirms that only the simplified hit-based firmware pieces have been combined. The central claim that an FPGA-based BGN-ST trigger can recover a large fraction of lost luminosity therefore requires a missing full-system measurement.","section":"Sections 4 and 5"},{"comment":"All reported accuracies are computed on simulated samples with a fixed 50% signal-to-background ratio, with no statistical uncertainties. The only realistic S/B study (0.1%) is for D0, where efficiency and purity are 23.2% and 2.3%; for beauty decays, Section 3.2 states \"purity and efficiency is currently under investigation.\" The 97.38% beauty accuracy is thus not yet connected to a realistic trigger performance metric (efficiency/purity at the expected S/B of ~0.05%).","section":"Sections 3.1 and 3.2"},{"comment":"The resource utilization quoted for the 8.82 us edge-classifier (194K LUT, 214K FF, 406 BRAM, 488 DSP) is for the Alveo U280; the paper does not provide resource projections for the smaller FELIX-712, nor a latency breakdown for the complete track-based pipeline. Without these, it is unclear whether pipelining the remaining stages can fit the 10 us target.","section":"Section 4"}],"minor_comments":[{"comment":"The author list and affiliations contain numerous character-encoding artifacts (e.g., \"/u1D44E\", \"/u1D450\") that must be fixed.","section":"Throughout"},{"comment":"There is a typo: \"modesl\" should be \"models\".","section":"Section 2"},{"comment":"The text contains \"BGS-ST\" where \"BGN-ST\" is meant.","section":"Section 4"},{"comment":"The parameter count \"363.170\" uses a period as the thousands separator, which is inconsistent with the other entries and should be \"363,170\" or \"363170\".","section":"Table 1"},{"comment":"The first sentence of the abstract is repeated verbatim in the main text; consider shortening the abstract to avoid redundancy.","section":"Abstract and Section 1"},{"comment":"The description of the simulated training data is brief; adding a reference or a sentence with the simulation parameters would improve reproducibility.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"This is a proceedings paper rather than a full technical article. The main technical concern is that the headline benefit is argued from sub-component latencies and simulated accuracy; the authors themselves are transparent about this (Section 5). In my view the paper is suitable for acceptance only after either (a) a full-system latency and resource measurement on FELIX-712, or (b) a revision that strictly limits claims to the demonstrated pieces and clearly labels the end-to-end target as projected. No concerns about the citation pattern or novelty."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a proceedings-style status report, and it reads like one. The genuinely new items are the beauty-decay accuracy (BGN-ST 97.38% on simulated events), the D0 improvement to 90.22% from data augmentation plus an adjacency-matrix loss, the comparisons against GarNet and GAT, and the first FPGA latency/resource numbers. The BGN-ST architecture itself was published before; this paper extends it.\n\nWhat it does well: the pipeline is clearly described, the baseline comparisons are useful, and the authors are upfront about what is not yet done. Section 5 states explicitly that full-system testing is expected by the end of the year. That honesty is worth something in an era of overclaimed trigger results.\n\nThe soft spots are real but not disqualifying for a status report. The stress-test note is correct that the 8.82 us edge-classifier latency on an Alveo U280 leaves little room in the 10 us target once you account for the FELIX-712 being about half the size. And the 9.2 us and 505 ns figures belong to the simpler hit-based GarNet, not the track-based BGN-ST that produces the headline accuracy. So the central claim—that this FPGA GNN trigger can recover a large fraction of the lost 90% luminosity—is plausible but not yet demonstrated end-to-end. The accuracy numbers also carry no error bars and were computed at a fixed 50% signal-to-background ratio; the efficiency/purity numbers are only for D0. No code or data are released, so independent checks are not possible. These are normal limitations for a proceedings contribution, but they limit how much weight the numbers can carry.\n\nI think the reader's conditional verdict is the right one. This is a useful snapshot of an ongoing DOE R&D project, and it will primarily interest people working on real-time ML in HEP triggers and DAQ. It deserves a serious referee if submitted to a peer-reviewed venue; the referee should push for the end-to-end FELIX-712 demonstration before the trigger claim is taken as fact. For your own work, it is citable as a project status reference, not as a demonstrated trigger solution.","headline":"An honest, useful status report on an FPGA GNN trigger for sPHENIX beauty decays; the new accuracy numbers are plausible, but the 10 microsecond end-to-end trigger claim is not yet demonstrated.","tokens_in":6307,"tokens_out":2615,"would_cite":true,"duration_ms":25960,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An FPGA-embedded graph neural network trigger can identify beauty decays at sPHENIX in microseconds, recovering events from the 90% of luminosity currently discarded.","keywords":["FPGA trigger","graph neural network","BGN-ST","beauty decays","streaming readout","real-time machine learning","hls4ml","sPHENIX"],"falsifier":"Run the full firmware chain on a FELIX-712 board with real 3 MHz p+p data and measure the latency from TPC buffer arrival to trigger output; if the end-to-end time exceeds roughly 30 microseconds, the buffer overflows and the claimed access to the discarded 90% of luminosity is lost. A second decisive check is to compare the beauty-enrichment of FPGA-triggered events against randomly triggered events offline; if no significant enrichment appears, the simulated accuracy does not translate to real data.","tokens_in":4889,"feed_emoji":"⚛️","tokens_out":7004,"duration_ms":64458,"temperature":0.7,"pith_summary":"The paper argues that machine-learning models embedded on FPGA hardware can do real-time track reconstruction and event filtering at high-rate collider detectors, and it demonstrates the pieces of such a system for the sPHENIX experiment. The specific claim is that an attention-based graph network, applied to tracks built from silicon-detector hits within the TPC's roughly 30-microsecond buffer, can pick out rare beauty-hadron decays from the 90% of sPHENIX luminosity that the current calorimeter trigger discards. On simulated data the model reaches 97.38% accuracy for beauty decays, and the track-based approach clearly outperforms hit-based and edge-based graph alternatives. If the full pipeline meets its 10-microsecond latency target, sPHENIX would gain access to a large sample of low-momentum heavy-flavor events with minimal extra resources, and the same approach could be reused for the future Electron-Ion Collider.","feed_headline":"FPGA neural net can recover 90% of sPHENIX luminosity","feed_subtitle":"A 10-microsecond track-based trigger spots beauty decays at 97.38% accuracy, unlocking data now discarded","key_machinery":"The central object is the Bipartite Graph Network with Set Transformers (BGN-ST), whose building block is called SEBA (Set Attention and Bipartite Aggregation). It represents silicon-detector hits as a graph, reconstructs tracks as edges, then uses attention to model track-to-track, track-to-global, and global-to-track interactions, iteratively assigning tracks to vertices and flagging displaced beauty-decay vertices. Around it, the hardware pipeline converts the trained model into FPGA firmware through manual C++ rewriting with the FlowGNN architecture and through automated translation with hls4ml; the 10-microsecond target matches the time available before the TPC buffer is overwritten.","core_discovery":"The paper's central discovery is that reconstructing tracks before the trigger decision, rather than classifying raw hits, gives a large accuracy boost for real-time heavy-flavor selection at sPHENIX. The BGN-ST model, a bipartite graph network with set transformers, uses 37 track features including cluster positions, inter-cluster edge lengths, angles, and track radius (momentum) to assign tracks to vertices and identify displaced decay topologies. In simulated beauty decays it reaches 97.38% accuracy, compared with 90.57% for a hit-based GarNet and 91.57% for a graph-attention model; in $D^0$ decays, adding track-radius (momentum) estimation improves trigger accuracy by 13.43 percentage points. The authors further report FPGA implementations: an edge-candidate classifier at 8.82 microseconds on a large Alveo U280 board, a simplified hit-based GarNet end-to-end at 9.2 microseconds, and a 505-nanosecond fully pipelined hls4ml version, with full-system testing scheduled by the end of the year.","pith_inferences":["The reported 97.38% accuracy was measured on 50%-signal/background simulated samples; real beauty production at RHIC is about 0.05%, so the metric that will decide practical value is efficiency and purity at realistic backgrounds, which the paper says is still under investigation.","If the full attention-based model cannot fit the latency budget on FELIX-712, a two-stage trigger could use the 505-nanosecond hit-based version as a prefilter and run BGN-ST only on surviving events, preserving much of the accuracy gain.","A working FPGA trigger of this type would make streaming readout viable for other experiments and could shift detector-design priorities toward keeping more data rather than rejecting it early.","The robustness test with exaggerated noise (65 extra hits, $10^{-7}$ noise versus the expected $10^{-9}$) suggests the hit-based front end degrades only about two percentage points; a similar stress test on track-based BGN-ST would indicate how much alignment uncertainty affects real-time performance."],"forward_implications":["If the 10-microsecond target is met, sPHENIX can save beauty-enriched data from the roughly 90% of delivered luminosity currently not recorded.","Track-based triggering raises heavy-flavor selection accuracy by 5–7 percentage points over hit-based or edge-based graph models, with the largest gain (13.43 points in $D^0$ studies) coming from estimating track momentum via curvature.","The $D^0$ trigger would give a 2.3-fold efficiency improvement over the current sPHENIX standard and a 23-fold rate improvement over random selection for tagging purity at 99% background rejection.","The same real-time workflow is being ported to an electron tagger for deep inelastic scattering at the future Electron-Ion Collider."],"supporting_citations":[{"why":"Supplies the graph convolutional network used to classify edge candidates in track construction.","marker":"[1]"},{"why":"Defines the BGN-ST architecture and earlier accuracy baselines that this paper extends and hardware-implements.","marker":"[2]"},{"why":"Provides the distance-weighted GarNet model used as the hit-based comparison and as the first end-to-end FPGA flow.","marker":"[3]"},{"why":"Supplies the graph attention network baseline compared for trigger detection from edge candidates.","marker":"[4]"},{"why":"Provides the FlowGNN dataflow architecture used to translate the hit-based model into FPGA firmware.","marker":"[5]"},{"why":"Supplies the hls4ml toolkit used for automated FPGA translation, including the 505-nanosecond pipelined implementation.","marker":"[6]"}],"fun_headline_variants":["Track-based AI trigger hits 97.38% accuracy at sPHENIX","FPGA GNN trigger: 505 ns, 97.38% beauty selection","Track-based trigger recovers discarded sPHENIX luminosity","Real-time GNN on FPGA selects beauty decays at 3 MHz","AI trigger outperforms hit-based by 6.8% at sPHENIX"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the complete track-based BGN-ST pipeline can finish within about ten microseconds on the FELIX-712 board, since the evidence shown covers only sub-components and simpler models, some tested on a larger Alveo U280 board.","fun_headline_variants_meta":{"raw":{"variants":["Track-based AI trigger hits 97.38% accuracy at sPHENIX","FPGA GNN trigger: 505 ns, 97.38% beauty selection","Track-based trigger recovers discarded sPHENIX luminosity","Real-time GNN on FPGA selects beauty decays at 3 MHz","AI trigger outperforms hit-based by 6.8% at sPHENIX"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000427,"raw_usage":{"total_tokens":2205,"prompt_tokens":982,"completion_tokens":1223,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":598,"completion_tokens_details":{"reasoning_tokens":1123}},"tokens_in":598,"tokens_out":1223,"duration_ms":8436,"temperature":1.0,"reasoning_tokens":1123,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:24:26.884655+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the full firmware chain on a FELIX-712 board with real 3 MHz p+p data and measure the latency from TPC buffer arrival to trigger output; if the end-to-end time exceeds roughly 30 microseconds, the buffer overflows and the claimed access to the discarded 90% of luminosity is lost. A second decisive check is to compare the beauty-enrichment of FPGA-triggered events against randomly triggered events offline; if no significant enrichment appears, the simulated accuracy does not translate to real data.","supporting_citations":[],"review_version":1}