{"id":"91bdf3c6-8312-4cdf-82da-0e376d6ceb6b","arxiv_id":"2607.13791","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A CNN likelihood-map plus (q,v) Bayes tracker and a motion-compensated matched-filter classical pipeline both estimate betatron tune from Schottky spectra at sub-millisecond latencies; the deep-learning estimator is most accurate at low SNR.","lead":"This paper develops two ways to measure the betatron tune in proton synchrotrons from noisy Schottky spectra: a classical signal-processing pipeline and a deep-learning estimator with a Bayesian tracker. On synthetic benchmarks both meet real-time latency; the deep-learning version is most accurate in low-SNR conditions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Synthetic-only benchmark cannot support the central low-SNR advantage until validated on independent spectra or real ramp data with a known tune reference.","rationale":"The paper is internally careful: the shared preprocessing is fixed, baselines are latency-compensated in a favorable way, ablations demonstrate that the tracker and motion compensation are load-bearing, and the real-beam data show end-to-end operation without divergence. However, the central claim that the deep-learning estimator is the most accurate across the operating range is validated only on a synthetic benchmark generated by the authors' own simulator, and the paper itself states that the real-beam test does not exercise the regimes where the estimators differ. This is the same load-bearing weakness the reader identified, and I agree with it. Because the concern is about external validity rather than an internal inconsistency, and because the reader already assigned CONDITIONAL, my read does not change the verdict. The proposed real-ramp test with an independent tune reference would settle whether the concern lands.","tokens_in":36131,"tokens_out":5723,"duration_ms":64471,"concrete_test":"Record Schottky spectra during a real SAPT acceleration ramp while obtaining an independent tune reference (e.g., a low-amplitude swept excitation or PLL on the BPM). Feed the frozen deep-learning and classical estimators the recorded frames without retraining. If the deep-learning estimator's MAE at ramp SNR at or below -10 dB does not remain below CNN+KF and the classical estimator by roughly the margins in Table 2, the simulator-fidelity concern lands. If the margins reproduce on independent real ramp data, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central accuracy comparison (Tables 2 and 3, Section 3.3) rests entirely on spectra from the authors' own Schottky simulator. Section 3.1 states \"no new Schottky-signal generator is introduced for this work [10]\", meaning the same simulator family is used for training the CNN, for setting classical-estimator parameters, and for evaluating both. Section 4 concedes that the real-beam acquisitions \"do not exercise the dynamic-tune, deep-low-SNR, or narrow-band regimes in which the two estimators are claimed to differ.\" If the simulator's noise model is statistically simpler than the operational chain—for example, white Gaussian noise plus calibrated sparse lines, without time-correlated pickup, revolution-frequency jitter, or slowly varying interference—then the deep-learning estimator's low-SNR advantage and the classical estimator's motion-compensated pooling may reflect matching the simulator's noise statistics rather than robust physical sideband recognition. In particular, the learned likelihood maps are trained on frames drawn from this generator, so the benchmark cannot fully distinguish \"recognizes Schottky sidebands\" from \"recognizes this simulator's spectra.\" This is the load-bearing premise behind the headline accuracy claim, and it is acknowledged in the paper's own limitation statement but not yet tested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents two betatron-tune estimators fed by a common folded-spectrum front end: a classical motion-compensated coherent-EMA estimator with a matched-filter bank and gated centroid, and a deep-learning estimator combining a convolutional likelihood-map network with a discrete (q,v) Bayes tracker. On a synthetic dynamic-tune benchmark spanning three trajectory speeds and five SNR levels, the deep-learning estimator attains the lowest aggregate MAE (geometric mean 0.187×10^-3 vs 0.243×10^-3 for CNN+KF and 0.418×10^-3 for the classical estimator) and the lowest catastrophic-miss rate; the classical estimator outperforms the latency-compensated T-PD baseline while requiring no training data or GPU. Robustness studies cover beam-loss recovery, narrow-band interference, out-of-distribution SNR, ablations, latency, input-length scaling, and acquisition resolution. A preliminary real-beam study at SAPT shows end-to-end operation and cross-estimator consistency on near-stationary coasting-beam acquisitions.","tokens_in":36441,"tokens_out":3927,"duration_ms":41874,"significance":"If the synthetic simulator faithfully represents the operational SAPT regimes, the contribution is significant: the likelihood-map-plus-Bayes-tracker architecture provides a large accuracy and tail-risk improvement in the low-SNR, dynamic-tune regime that is most relevant to compact medical synchrotrons, while the classical estimator offers a practical CPU-only, training-free alternative. The paper is careful in several respects: all estimators receive identical input streams; the deep-learning results are cross-validated over five folds; ablations decompose the tracker and CNN contributions; latency is measured on deployed hardware; and the authors explicitly acknowledge the synthetic-only evidence for the headline regime. The main limitation is that the central dynamic-tune and low-SNR claims are not validated against independent spectra or real ramp data with a known tune reference, and some tracker motion constants are calibrated using the benchmark trajectories themselves, which limits the strength of the claims as currently stated.","major_comments":[{"comment":"The central accuracy claim (Tables 2-3) rests entirely on spectra from the authors' own Schottky simulator, referenced to [10] with \"no new Schottky-signal generator introduced.\" The paper's own limitation statement (Section 4) concedes that the real-beam acquisitions \"do not exercise the dynamic-tune, deep-low-SNR, or narrow-band regimes in which the two estimators are claimed to differ.\" As a result, the benchmark cannot distinguish recognition of betatron sidebands from recognition of this simulator's spectral statistics. The claims should either be explicitly scoped to synthetic spectra throughout, or the authors should provide independent validation (e.g., an external simulator, measured spectra with an independent tune reference, or a hardware-in-the-loop test) before the operational generalization is asserted.","section":"Sections 3.1, 3.3, 3.6, 4"},{"comment":"The tracker motion constants are calibrated to the benchmark trajectories: sigma_v is set \"so that 4 sigma_v covers the largest per-frame velocity change measured on the evaluated trajectories,\" and v_max is chosen to \"cover... the fastest ramp trajectories evaluated.\" Evaluating the estimator on the same trajectory family used to set its motion-model constants is a circularity risk for the reported margin over baselines. Please provide a sensitivity analysis over sigma_v and v_max, and/or evaluate on held-out trajectory ensembles with velocity statistics outside the calibration range, to show that the deep-learning advantage is not an artefact of matching the benchmark's motion envelope.","section":"Section 2.5 (Eq. 2.20) and Section 3.1"},{"comment":"The drifting-clutter robustness result, which is cited as a key differentiator, is based on a single representative 600-frame energy ramp with three interference lines; the five-cell geometric mean is over five SNR levels of that one ramp, not over independent ramp realizations. Because the frames are strongly autocorrelated, the reported per-cell statistics have very low effective sample size. The paper should report results over multiple random ramp realizations (with randomized line positions/amplitudes) or provide a statistical justification for treating the single ramp as sufficient. The same concern applies to the stationary-clutter stream.","section":"Section 3.4.2 (Tables 5-6)"}],"minor_comments":[{"comment":"The no-tracker deep-learning readout is described inconsistently: Section 3.2 uses a \"peak-local soft-argmax,\" while Table 12 labels the row \"DL, no tracker\" as a \"whole-map soft-argmax.\" Please clarify which readout is actually used in each table and ensure the metric is defined consistently.","section":"Sections 3.2 and 3.5.4"},{"comment":"Panel (b) is captioned \"fast, 15 dB\" but the text and the surrounding discussion refer to the -15 dB condition. Please correct the sign or the label.","section":"Figure 7 caption"},{"comment":"Minor typos: \"kernel-size-71D convolution\" should read \"kernel-size-7 1D convolution\" (two typos in one phrase), and τ=0.5 is described as \"bin\" where the temperature is dimensionless. These should be cleaned up before publication.","section":"Section 2.5"},{"comment":"The latency-compensated T-PD comparison selects N* per trajectory shape to minimize MAE on the same benchmark used for the headline comparison. This is a permissive treatment of the baseline; the paper should state whether the reported N*=3 also transfers to the robustness studies, or whether it was re-selected there.","section":"Section 3.3 / T-PD baseline"},{"comment":"For the CNN+KF baseline, the reported mean posterior sigma (0.264×10^-3) is far below the realized within-acquisition std (6.132×10^-3), indicating severe miscalibration. The text notes this only implicitly. It would be helpful to add an explicit caution that the CNN+KF uncertainty is not calibrated on this real-beam set, especially since the proposed deep-learning estimator is compared against it.","section":"Table 14"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is technically careful and the synthetic evaluation is internally consistent, but the main operational claim is supported only by the authors' own simulator in the regimes that matter. The limitations are stated honestly in Section 4, which is to the authors' credit, but the abstract and conclusion currently present the synthetic results as establishing the operational advantage. A major revision that (a) reframes the claims to the synthetic scope or adds independent validation, and (b) addresses the motion-constant calibration circularity and the single-ramp clutter statistics, would make this suitable for publication. I do not see a load-bearing internal error that would require rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: two new estimators, both described in enough detail to be built from the paper, and the synthetic benchmark is run carefully. But the claim that matters — the deep-learning estimator's low-SNR edge — is demonstrated only on spectra from the authors' own simulator, and the real-beam data explicitly does not cover the regimes where the two estimators differ. Treat the headline numbers as provisional.\n\nThe paper does a lot well. The classical motion-compensated coherent EMA with matched-filter bank and MAD-gated centroid is a coherent, inventive design; the deep-learning pipeline (likelihood-map CNN with FFT global convolutions plus a discrete (q,v) Bayes tracker) is a sensible separation of spectral recognition from temporal inference. The ablations are illuminating: the tracker decomposition shows the (q,v) posterior is what gives the deep estimator its edge, and the classical EMA ablation shows both pooling and motion compensation matter. Cross-fold consistency, latency measurements, and the honest scope statement in Section 4 all speak for the authors.\n\nThe soft spots are real, and they are load-bearing. The dynamic-tune benchmark is generated by the same simulator family used to train the CNN and to set the classical estimator's parameters. That cannot fully separate 'recognizes Schottky sidebands' from 'recognizes this simulator's noise.' The tracker's velocity envelope and process noise are explicitly calibrated to the benchmark trajectories, which is a form of leakage into the evaluation. The real-beam section is a feasibility and cross-estimator agreement check, not a validation of the claimed operating regimes, and the authors say so. No code or data are released, so the numbers are currently not independently checkable.\n\nNone of this makes the paper incoherent. The central argument holds up as far as it goes: the proposed methods are new, the experimental design is internally consistent, and the limitations are stated. What is missing is independent evidence. If the low-SNR advantage survives on spectra from a different source, or on ramp data with a known tune reference, this becomes a solid contribution to accelerator instrumentation.\n\nWho should read it: anyone working on non-invasive tune measurement in compact medical synchrotrons or similar low-SNR Schottky diagnostics. It deserves a serious referee: the method detail is enough to reproduce, and the claims, if independently confirmed, matter. I would send it to peer review, with an explicit request that the synthetic-only validation be addressed or clearly scoped.","headline":"Honest, carefully built paper with two genuinely new estimators, but the headline low-SNR advantage rests entirely on a synthetic benchmark from the authors' own simulator, and real beam only shows feasibility.","tokens_in":36914,"tokens_out":2923,"would_cite":false,"duration_ms":53307,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A CNN likelihood map plus a discrete (q,v) Bayes tracker gives the most accurate betatron-tune readout from Schottky spectra at low SNR, with a training-free classical counterpart close behind.","keywords":["Schottky spectra","betatron tune","slow extraction","medical proton synchrotron","Bayesian tracking","convolutional neural network","motion compensation","low-SNR diagnostics"],"falsifier":"Acquire Schottky spectra during a real accelerating ramp or slow-extraction cycle at the target facility with an independently known tune (for instance, a kicker-excitation reference or a dedicated tune meter), then measure the geometric-mean absolute error of the deep-learning estimator versus the classical estimator and the prior CNN+Kalman baseline at effective SNR near -15 dB. If the deep-learning estimator's error advantage over the classical and Kalman baselines fails to appear, or reverses, on real data, the simulator-fidelity premise — and with it the central comparison — collapses.","tokens_in":36031,"feed_emoji":"⚛️","tokens_out":5479,"duration_ms":54125,"temperature":0.7,"pith_summary":"The paper sets out to make passive betatron-tune measurement from Schottky spectra reliable in the difficult conditions of compact medical proton synchrotrons: short acquisition windows, low signal-to-noise ratio, and residual narrow-band interference. It develops two estimators that share one spectral front-end but accumulate temporal context in different ways: a classical motion-compensated coherent moving average feeding a matched-filter bank and gated centroid, and a deep-learning pipeline in which a convolutional network converts each spectrum into a tune-likelihood map that a discrete (q,v) Bayesian filter propagates frame to frame. On a synthetic dynamic-tune benchmark the deep-learning estimator has the lowest mean absolute error (geometric mean 0.187e-3 tune units) and by far the lowest rate of catastrophic misses (0.14 percent), while the classical estimator beats a latency-compensated published baseline and needs no training data or GPU. Both run end-to-end on real beam data from the target facility without retraining, and median per-frame latency stays below 1 ms.","feed_headline":"CNN-plus-Bayes tracker wins low-SNR Schottky tune readout","feed_subtitle":"Lowest mean error 0.187e-3 tune units; a training-free CPU-only classical rival still beats the prior baseline.","key_machinery":"The shared front-end is the folded PSD on a uniform 1024-bin tune grid. The classical estimator's key mechanism is the motion-compensated coherent exponential moving average, which shifts the pooled spectrum by a velocity estimate before averaging, with a velocity-adaptive decay, then detects the sideband with a multi-width Gaussian matched-filter bank and refines it with a median/MAD-gated centroid. The deep-learning estimator's key mechanism is the combination of a likelihood-map CNN — using FFT-evaluated global convolutions with an implicit damped-oscillator kernel, giving full-axis receptive field at O(L log L) cost — and a discrete (q,v) Bayes tracker that propagates a log-posterior ove","core_discovery":"The central claim is that betatron-tune readout from low-SNR Schottky spectra is best served not by choosing between classical and learned per-frame estimators, but by pairing a fixed spectral front-end with a temporal representation matched to the estimator's strengths. The deep-learning estimator's advantage comes from feeding a full tune-grid likelihood map, not a scalar, into a discrete two-dimensional (q,v) Bayes tracker: the multi-modal map keeps a wrong per-frame peak recoverable, and the velocity-aware posterior rejects clutter and noise while reporting per-frame uncertainty. The classical estimator achieves competitive accuracy at moderate SNR with a velocity-shifted coherent EMA th","pith_inferences":["A natural next test is to collect real ramping-cycle Schottky data with an independent tune reference (e.g., a kicker-based measurement on a sacrificial cycle) and compare estimators at -10 to -20 dB; if the deep-learning advantage does not survive the real noise structure, the paper's central comparison rests on simulator fidelity.","The length-portability of the CNN suggests the same trained network could be applied to spectra with different harmonic grids, or even to other frequency-domain diagnostics such as chromaticity from Schottky spectra, provided the coordinate and weight channels are scaled accordingly.","The observed complementarity hints at a practical fusion: run both estimators and treat large disagreement as a flag for anomalous spectra (beam loss, interference, or model mismatch), rather than committing to a single estimator.","One could ablate further by training on intentionally mismatched noise models (colored, non-Gaussian) to see whether the likelihood-map tracker retains robustness, which would separate simulator-fidelity effects from architectural benefits."],"forward_implications":["If the simulator fidelity holds, the deep-learning estimator could replace kicker-excitation tune measurement with passive low-SNR Schottky readout during slow extraction, avoiding beam disturbance.","The discrete tracker's per-frame posterior standard deviation provides a ready-made uncertainty channel for gating downstream systems such as spill-quality feedback — something the classical estimator does not provide.","The classical estimator's tune-unit parametrization and training-free operation mean it can be deployed to other synchrotrons with minimal recalibration, making it a safe default when GPU or training infrastructure is absent.","The operating-regime decision matrix gives a practical rule: use the deep-learning estimator whenever SNR is tight or unpredictable, and the classical estimator at moderate SNR on benign, slowly varying tunes.","Both estimators' sub-millisecond latency (0.12 ms on CPU, 0.40 ms on GPU) supports real-time, closed-loop tune measurement at millisecond frame cadences."],"fun_headline_variants":["Schottky tune via CNN-Bayes beats baselines, classical rival stays CPU-only","Bayesian CNN tracker wins low-SNR Schottky tune; classical rival needs no GPU","Deep-learned tune map beats baselines; classical CPU-only rival shines","Two-way tune readout: CNN-Bayes leads, classical no-training CPU rival holds","Schottky tune: deep map wins, classical CPU-only rival beats baseline"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the synthetic Schottky simulator used to train and benchmark both estimators faithfully represents the target facility's real spectra in exactly the regimes where the claims are made — the paper's own real-beam acquisitions are near-stationary and do not exercise the low-SNR, dynamic-tune, or narrow-band conditions that separate the estimators.","fun_headline_variants_meta":{"raw":{"variants":["Schottky tune via CNN-Bayes beats baselines, classical rival stays CPU-only","Bayesian CNN tracker wins low-SNR Schottky tune; classical rival needs no GPU","Deep-learned tune map beats baselines; classical CPU-only rival shines","Two-way tune readout: CNN-Bayes leads, classical no-training CPU rival holds","Schottky tune: deep map wins, classical CPU-only rival beats baseline"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000927,"raw_usage":{"total_tokens":3827,"prompt_tokens":782,"completion_tokens":3045,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":2936}},"tokens_in":526,"tokens_out":3045,"duration_ms":18482,"temperature":1.0,"reasoning_tokens":2936,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T03:40:27.299126+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Acquire Schottky spectra during a real accelerating ramp or slow-extraction cycle at the target facility with an independently known tune (for instance, a kicker-excitation reference or a dedicated tune meter), then measure the geometric-mean absolute error of the deep-learning estimator versus the classical estimator and the prior CNN+Kalman baseline at effective SNR near -15 dB. If the deep-learning estimator's error advantage over the classical and Kalman baselines fails to appear, or reverses, on real data, the simulator-fidelity premise — and with it the central comparison — collapses.","supporting_citations":[],"review_version":1}