{"id":"ce581ee3-8623-4182-af99-ac8554187ffd","arxiv_id":"2411.13596","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A GNN-based end-to-end track finder for the Belle II drift chamber reconstructs displaced tracks at 85.4% efficiency with a 2.5% fake rate, outperforming the baseline algorithm at 52.2%.","lead":"A graph neural network now reconstructs charged particle tracks in Belle II's drift chamber directly from detector hits, predicting track count and kinematics in a single pass. On simulated displaced decays it finds and fits 85.4% of tracks with a 2.5% fake rate, versus 52.2% for the current Belle II algorithm, which matters for searches for long-lived new particles.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 2.5% fake rate for dark-Higgs tracks is not shown to be stable against the un-tuned GENFIT2 initial covariance matrix (fixed to 0.1 without a scan), leaving a self-acknowledged hole in the claimed comparison.","rationale":"The reader's weakest assumption is simulation fidelity; that is a legitimate limitation but not an internal gap, because the paper explicitly frames the results as using a realistic full detector simulation, so the simulation-based claim is supported as stated. The more load-bearing internal issue is the arbitrary initial covariance handed to the track fitter, because the headline fake rate is a post-fit quantity and the fitter demonstrably rejects many CAT tracks (pre-fit fake rate 5.12% vs post-fit 2.12% in the barrel). The paper acknowledges the covariance was not optimized and checks only efficiency, not fake rate. A simple scan can settle whether the 2.5% fake rate is a property of the GNN or an artifact of the chosen covariance scale. This is also a concrete instance of the systematic uncertainties the reader's conditional verdict asks for, so the verdict stays conditional. The covariance scan is cheap because the model and samples already exist; it should be part of the response to the review.","tokens_in":37012,"tokens_out":20054,"duration_ms":201119,"concrete_test":"Run the CAT Finder on the statistically independent dark Higgs sample (Section 7.1.3) for a log-spaced grid of initial covariance scales, e.g. multiply the 0.1 entries by {0.1, 0.3, 1, 3, 10}, keeping all other settings fixed. Recompute the combined track finding and fitting charge efficiency, fake rate, and clone rate (Eqs. 5, 7, 8) and compare to Table 4. If the fake rate moves by more than 1 percentage point around 2.5%, or the charge efficiency shifts by more than 0.5%, the headline comparison is not robust to this unoptimized parameter; the paper would then need to report the systematic band or optimize the covariance on a held-out sample.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result (Section 7.1.3, Table 4) is a combined track finding and fitting charge efficiency and fake rate: 85.4% and 2.5% for the CAT Finder versus 52.2% and 4.1% for the baseline. The fitting step uses GENFIT2 with a DAF, whose hit weighting depends on the initial covariance matrix. Section 6.3 states: \"All initial covariance matrix entries for the CAT Finder are set to 0.1. We observe that the track fitting time depends on the initial parameters of the covariance matrix. However, the impact on the track finding efficiency is negligible and we leave performance optimization of this to future work.\" Only the efficiency is checked; the fake rate and charge efficiency are not scanned. The DAF's annealing thresholds are scaled by the covariance; a too-small covariance rejects many hits (and fakes), a too-large covariance accepts ambiguous hits. Because the CAT Finder's post-fit fake rate in the barrel (2.12%, Table 4) is much lower than its pre-fit fake rate (5.12%), the fitter is doing substantial fake rejection. That rejection could be an artifact of the arbitrary 0.1 scale rather than a property of the GNN. If a different scale raises the post-fit fake rate toward the baseline level, the claim \"with a fake rate of 2.5%\" loses its quantitative force. The paper itself flags this as unsupported, so the missing scan is a self-acknowledged gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents CAT Finder, a graph neural network (GNN) based end-to-end track finder for the Belle II central drift chamber. Using wire hits as node inputs, the network predicts an unknown number of tracks via object condensation, estimates each track's momentum, charge, and starting position, and clusters hits for later fitting by GENFIT2 with a deterministic annealing filter. The study is performed on GEANT4-based full detector simulation with beam backgrounds and detector conditions overlaid from recorded high-luminosity data. The central result is a combined track finding and fitting charge efficiency of 85.4% per track with a 2.5% fake rate for dark-Higgs-like displaced decays, compared with 52.2% and 4.1% for the Belle II baseline CDC finder. Additional results cover prompt tracks, radiative muon pairs, and K0S decays, along with robustness checks against background conditions and a discussion of lessons learned during training.","tokens_in":37243,"tokens_out":6133,"duration_ms":69093,"significance":"If the reported performance holds, this is a valuable demonstration of an end-to-end machine-learned track finder in a realistic drift chamber environment, with substantial gains over the existing algorithm for displaced decays and for the forward/backward endcap regions. The paper's strengths include statistically independent evaluation samples (about one million events total), realistic simulation including data-overlay beam backgrounds, a documented check that removing physical training samples avoids a specific overfitting mode, an explicit robustness study across background conditions, and public code for reproduction. The central claim is an empirical benchmark rather than a derivation, so there is no circularity concern; the main risk is whether the headline fake rate is robust to an arbitrary track-fitter initialization, as detailed in the major comments.","major_comments":[{"comment":"The combined track finding and fitting fake rate of 2.5% (barrel 2.12% after fitting) is a headline claim, but it depends on the GENFIT2 deterministic annealing filter, whose hit-acceptance thresholds scale with the initial covariance matrix. All initial covariance entries for CAT Finder are set to 0.1, and the manuscript states only that the impact on track finding efficiency is negligible and that performance optimization is left to future work. The post-fit fake rate in the barrel drops from 5.12% (track finding) to 2.12% (after fitting), so the fitter is performing substantial fake rejection. A different diagonal covariance scale could plausibly change the DAF's hit rejection and thus the post-fit fake rate and charge efficiency. The paper should either scan the initial covariance scale (e.g., over factors from 0.1 to 10) and report efficiency, fake rate, and charge efficiency, or provide a principled initialization (for example, a covariance derived from the GNN parameter predictions) and show that the headline 2.5% fake rate is stable under this choice.","section":"Section 6.3, Table 4"}],"minor_comments":[{"comment":"The main prompt-track comparison removes curling tracks from the evaluation, so the claim of significantly better forward/backward performance is specific to the non-curling subset. This restriction should be stated more prominently in Section 7.1.1 and in the captions of Fig. 7 and Table 3, since the integrated prompt numbers are not directly comparable to a full inclusive sample.","section":"Section 7.1.1, Fig. 7, Table 3"},{"comment":"The caption of Table D1 states that the muon-pair results are for low beam background, while Section 7.1.2 and Fig. D5 describe these samples as using high data beam backgrounds; this inconsistency should be corrected.","section":"Appendix D, Table D1"},{"comment":"Please clarify the units and structure of the initial covariance matrix (diagonal versus full, units of momentum and position entries), and briefly state what the Baseline Finder uses for its initial covariance so the reader can judge the fairness of the comparison.","section":"Section 6.3, covariance text"},{"comment":"The relationship between the track charge efficiency and the wrong-charge rate could be stated more explicitly: a matched track with the wrong charge is counted in the wrong-charge rate but not in the charge efficiency, so the two metrics do not sum to unity; making this explicit would avoid confusion when reading the tables.","section":"Section 5, Eqs. (5) and (9)"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong, honest empirical study with reproducible code and a clearly scoped claim. The single major issue—robustness of the headline fake rate to the arbitrary initial covariance of the GENFIT2 fitter—is exactly the kind of gap that a careful referee should require to be closed before publication. It is fixable within the manuscript's scope by adding a covariance-scale scan or a simulation-derived initialization. I also note that the 'first end-to-end multi-track ML algorithm' claim is defensible for track finding, though the fitting step remains conventional; the authors might soften the wording slightly in the conclusion to avoid overstatement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. The CAT Finder is the first end-to-end, unknown-number multi-track reconstruction for a drift chamber in a realistic setting, and the displaced-track result is the real news: 85.4% vs 52.2% combined charge efficiency with a lower fake rate on dark-Higgs-like decays. The evaluation is more careful than most ML tracking papers: independent samples, statistical uncertainties, an explicit overfitting check, and a robustness study across beam backgrounds. They also make code available. Credit where due: the GNN components are not novel in isolation—object condensation and GravNet are established—but the integration with raw CDC hits, displaced vertices, and a full GENFIT2 fit chain is new and well executed.\n\nSoft spots, in proportion. First, the prompt comparison in the main text excludes curling tracks, which are a real part of the low-pt phase space. The paper does analyze curlers separately and is transparent about it, but the abstract's 'variety of event topologies' overstates what the headline prompt numbers cover. Second, the tβ, td, th working point was optimized on category 2 and Ks barrel, so the prompt barrel numbers in Table 3 are not fully independent. That does not touch the dark Higgs result, which is the main claim. Third, and the one a referee should push on: the GENFIT2 initial covariance is fixed to 0.1 with no scan. The authors say it does not affect finding efficiency, but the fitter also rejects fakes—the post-fit fake rate drops from 5.12% to 2.12% in the barrel. Without a covariance scan, we do not know how much of that fake-rate suppression is a property of the GNN and how much is an artifact of the arbitrary initial scale. The paper discloses this, but the headline 2.5% fake rate rests on it. That is a missing systematic, not a fatal flaw.\n\nOverall: solid, useful, and worth a serious referee. I would accept for peer review and ask for three things: a curler caveat in the abstract, a working-point independence check for the prompt barrel numbers, and a covariance scan to bound the fake-rate systematic. Recommendation: engage with it and send it to review.","headline":"A credible, well-evaluated GNN end-to-end tracker for the Belle II drift chamber; the displaced-track gain holds up, but the prompt-number caveats and an unscanned fitter covariance deserve referee attention.","tokens_in":37929,"tokens_out":2724,"would_cite":true,"duration_ms":26779,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["29.40.Gx","07.05.Mh"],"model":"deepseek-v4-flash","headline":"A graph neural network that reads Belle II drift-chamber hits directly reconstructs macroscopically displaced tracks at 85.4% combined efficiency with a 2.5% fake rate, far above the baseline algorithm's 52.2% and 4.1%.","keywords":["track finding","graph neural networks","object condensation","drift chamber","Belle II","displaced vertices","end-to-end reconstruction","particle tracking"],"falsifier":"Run the trained CAT Finder on real Belle II collision data for $K^0_S \\to \\pi^+\\pi^-$ decays, where the decay vertex is independently measured, and compare per-track finding-and-charge efficiency and fake rate in the barrel with the simulation's ~93% efficiency and ~5% fake rate; a large shortfall would show the simulated drift-time or wire-inefficiency model is not faithful for displaced tracks.","tokens_in":36750,"feed_emoji":"⚛️","tokens_out":8426,"duration_ms":74234,"temperature":0.7,"pith_summary":"The paper tries to establish that a single graph neural network, the CAT Finder, can replace the conventional track-finding stage of Belle II's drift chamber reconstruction: it consumes raw wire hits with no pre-filtering, predicts the unknown number of charged tracks together with their momenta, starting points, and charges, and then hands each predicted track's hits to a standard Kalman filter. If correct, this would give a generic way to do tracking where detectors measure drift time rather than direct three-dimensional positions, which is exactly the regime where the existing Legendre-transform baseline loses efficiency for tracks from decays displaced by centimeters to a meter. The paper reports that for dark-Higgs-like displaced decays the CAT Finder achieves a combined finding-and-fitting charge efficiency of 85.4% per track with a 2.5% fake rate, compared to 52.2% and 4.1% for the baseline, and that for prompt tracks it performs comparably in the barrel while substantially better in both endcaps. This matters for experimental particle physics because displaced-vertex signatures are a primary search route for long-lived new particles.","feed_headline":"Neural tracker hits 85.4% efficiency on displaced tracks at Belle II","feed_subtitle":"Beats the 52.2% baseline for displaced decays and lifts endcap tracking in full simulation with real beam backgrounds.","key_machinery":"The mechanism is a graph neural network whose nodes are individual wire hits, processed by four GravNet blocks, a distance-weighted graph layer that builds edges between nodes in a learned representation space and aggregates features along those edges. Object condensation is the central identity: a loss that makes each node predict a condensation coordinate in a learned low-dimensional cluster space and a confidence value $\\beta$, pulling hits of the same particle toward one high-$\\beta$ condensation point while repelling different particles; after inference, thresholds on $\\beta$ and on cluster-space distances ($t_\\beta$, $t_d$, $t_h$) select the condensation points that define the tracks and their parameter predictions, and hits within a radius $t_h$ of a condensation point form the track's hit set for the subsequent fitter. This learned cluster space is what carries the argument, because it converts the unknown-number-of-tracks problem into a clustering problem whose solution is read off geometrically.","core_discovery":"The central claim is that end-to-end multi-track reconstruction in a drift chamber is feasible and outperforms the conventional finder for the most challenging topologies: the CAT Finder simultaneously predicts the number of tracks, each track's three-momentum, starting point, and charge, and assigns the detector hits that belong to it, using object condensation in a learned cluster space. The evaluation is deliberately realistic: full GEANT4 detector simulation with digitized detector response, beam-background hits overlaid from actual collision data, and wire-efficiency maps from data-taking conditions. Against this backdrop the GNN finds and correctly charges 85.4% of tracks from displaced dark-Higgs-like decays with a 2.5% fake rate, versus 52.2% and 4.1% for the baseline; it is also significantly better in the endcaps for prompt tracks and comparable in the barrel. The paper presents this as the first end-to-end machine-learning multi-track reconstruction for a drift chamber operated in a realistic particle physics environment.","pith_inferences":["The same architecture should transfer to other drift chambers and gaseous tracking detectors that, like the Belle II CDC, lack direct spatial readout and must infer positions from drift time; the network learns the wire geometry, so only the input feature encoding would need to change.","The paper's observation that the network infers a vertex position even without nearby hits, using the companion track's direction, hints that a GNN of this kind could provide trigger-level vertexing for displaced signatures without a separate vertex-finding step.","The large gap between the CAT Finder's hit efficiency for curling tracks and the conventional fitter's success on them suggests that the reported 85.4% is limited by the fitter, not the finder; a fitter that accepts multi-curler hit sets could push efficiency higher.","The fidelity of the simulated wire inefficiency and cross-talk is untested on real displaced tracks, so the magnitudes of the reported gains should be re-measured on data before relying on them."],"forward_implications":["Long-lived-particle searches at Belle II (e.g., dark Higgs and inelastic dark matter) would see a large efficiency jump: for a benchmark dark Higgs with $m_h=1.5\\,\\text{GeV}$ and $c\\tau=21.5\\,\\text{cm}$, both tracks are reconstructed in 87.2% of events versus 44.9% with the baseline, at a lower fake rate.","Endcap tracking, historically the weakest region for the baseline finder, improves to near-barrel levels for both prompt and displaced tracks, increasing the usable acceptance for forward and backward physics.","Because the GNN outputs track parameters directly from the condensation point, the hit clustering and Kalman-filter steps can be dropped for real-time triggering, opening a route to FPGA-based end-to-end tracking.","The robustness test (training on low simulated background, evaluating on high background from data) shows only a few-percent efficiency loss without retraining, suggesting the model can ride out changing accelerator conditions."],"supporting_citations":[{"why":"Supplies the dark Higgs / inelastic dark matter model that defines the displaced-decay evaluation samples.","marker":"[3]"},{"why":"Provides the beam-background expectations used to simulate realistic CDC occupancy and overlay data backgrounds.","marker":"[4]"},{"why":"Provides the GravNet distance-weighted graph architecture used to learn the irregular CDC wire geometry.","marker":"[7]"},{"why":"Supplies the object condensation loss that lets the network predict an unknown number of tracks from hits.","marker":"[9]"},{"why":"Defines the baseline Belle II track finder against which the CAT Finder is compared.","marker":"[29]"},{"why":"The GEANT4 full detector simulation that produces the signal hits and material interactions used for training and evaluation.","marker":"[31]"},{"why":"The basf2 framework that turns simulated energy deposits into digitized CDC hits with detector response.","marker":"[32, 33]"},{"why":"GENFIT2, the Kalman-filter track fitter that converts CAT Finder hit sets into fitted tracks; the combined efficiency includes this step.","marker":"[38–41]"}],"fun_headline_variants":["GNN finds tracks Belle II missed, 85% vs 52%","Belle II GNN: 85% track find efficiency on displaced decays","GNN outperforms Belle II baseline: 85% vs 52% on displaced tracks","First end-to-end GNN tracker at Belle II finds 85% displaced"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results stand or fall on whether the GEANT4 simulation with digitized detector response, data-overlay beam backgrounds, and the corresponding wire-inefficiency maps faithfully reproduces the real Belle II CDC response for tracks that originate far from the interaction point.","fun_headline_variants_meta":{"raw":{"variants":["GNN finds tracks Belle II missed, 85% vs 52%","Belle II GNN: 85% track find efficiency on displaced decays","GNN outperforms Belle II baseline: 85% vs 52% on displaced tracks","First end-to-end GNN tracker at Belle II finds 85% displaced"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000944,"raw_usage":{"total_tokens":4042,"prompt_tokens":965,"completion_tokens":3077,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":581,"completion_tokens_details":{"reasoning_tokens":2991}},"tokens_in":581,"tokens_out":3077,"duration_ms":18096,"temperature":1.0,"reasoning_tokens":2991,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:04:18.185165+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained CAT Finder on real Belle II collision data for $K^0_S \\to \\pi^+\\pi^-$ decays, where the decay vertex is independently measured, and compare per-track finding-and-charge efficiency and fake rate in the barrel with the simulation's ~93% efficiency and ~5% fake rate; a large shortfall would show the simulated drift-time or wire-inefficiency model is not faithful for displaced tracks.","supporting_citations":[{"cited_title":"Duerr et al","cited_arxiv_id":null,"evidence_quote":"Supplies the dark Higgs / inelastic dark matter model that defines the displaced-decay evaluation samples."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the GravNet distance-weighted graph architecture used to learn the irregular CDC wire geometry."},{"cited_title":"Kieseler","cited_arxiv_id":null,"evidence_quote":"Supplies the object condensation loss that lets the network predict an unknown number of tracks from hits."},{"cited_title":"Bertacchi et al","cited_arxiv_id":null,"evidence_quote":"Defines the baseline Belle II track finder against which the CAT Finder is compared."},{"cited_title":"Agostinelli et al","cited_arxiv_id":null,"evidence_quote":"The GEANT4 full detector simulation that produces the signal hits and material interactions used for training and evaluation."}],"review_version":1}