{"id":"bcff27e4-ffc5-4d0b-9ae9-b6761beb6357","arxiv_id":"2607.02987","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Hybrid online-offline SMC with batch-averaged Kraus maps estimates time-dependent multiparameters from noisy continuous quantum measurements better than calibration on two superconducting-qubit datasets.","lead":"A hybrid sequential Monte Carlo method averages noisy quantum measurement batches into trajectories and tracks drifting multiparameters via approximated Kraus maps. It outperforms standard calibration on superconducting-qubit fluorescence and dispersive data, including detection of an uncalibrated jump.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Signal-reconstruction RMSE is not an independent ground-truth test of multiparameter accuracy; model incompleteness can produce lower RMSE with systematically wrong parameters.","rationale":"The Reader correctly flags the constant-within-batch assumption and the Γ–η degeneracy, but treats the reconstruction RMSE as confirmatory evidence of superior multiparameter accuracy. That is the more load-bearing soft spot: the validation metric is not independent of the model class being fitted. The paper already shows (Fig. 6, Appendix B) that the likelihood surface is degenerate and that unmodeled decoherence can be absorbed; therefore a lower RMSE does not logically entail that the estimated parameters are closer to truth. The methodological contribution and the modular SMC derivation remain solid, so the verdict stays CONDITIONAL rather than REJECT, but the experimental claims should be read more cautiously until an independent ground-truth test is supplied.","tokens_in":27050,"tokens_out":595,"duration_ms":6217,"concrete_test":"Generate synthetic fluorescence and dispersive trajectories from the exact SMEs of Sec. 4.2–4.3 with known, slowly drifting parameters plus an unmodeled T_{2} term; run the identical SMC pipeline (same Ns, Np, p, q) and compute both parameter RMSE versus ground truth and signal-reconstruction RMSE versus the static calibration. If signal RMSE improves while parameter RMSE remains large (or the recovered jump is an artifact of the misspecification), the experimental claim is unsupported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper’s central experimental claim (fluorescence RMSE 0.37 vs 0.75; recovery of a bias jump after batch 17) rests on the premise that lower reconstruction error of the batch-averaged voltage trajectory implies more accurate multiparameter estimates. That premise is insecure. The observation model (Eqs. 24/29 and the O(Δτ^{2}) map of Appendix A) is incomplete: for the dispersive case the authors themselves document a near-perfect Γ–η degeneracy (Fig. 6) and note that unmodeled T_{2} dephasing can be absorbed into the same product. Consequently a biased set of parameters can still generate a mean trajectory that matches the experimental average better than the laboratory calibration, simply because the SMC particles are free to trade off the free parameters against residual model error. The fluorescence comparison is likewise only a relative ranking against a static calibration that was never claimed to be simultaneous with the continuous-record data. Without an independent ground-truth check (or a controlled simulation in which the true parameters are known and the model is deliberately misspecified), the RMSE numbers and the reported jump cannot be taken as conclusive evidence that the hybrid estimator recovers the true time-dependent multiparameter trajectory.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper develops a hybrid online-offline sequential Monte Carlo (SMC) algorithm for time-dependent multiparameter estimation under continuous quantum measurement. It combines sequential importance sampling (with a static/zero-noise transition) as the main recursion, an SMC-sampler resampling step to combat degeneracy, and batch partitioning of noisy trajectories. Batch-averaged signals are evolved with an efficient O(Δτ^{2}) approximation to the trajectory-averaged Kraus map (derived in Appendix A). Hyperparameters are selected via independent numerical simulations (Appendix B). The method is demonstrated on two superconducting-qubit datasets (fluorescence and dispersive z-measurement). Validation uses signal reconstruction RMSE of batch-averaged voltage traces against laboratory calibrations; the authors report lower RMSE than calibration for fluorescence and recovery of an unreported jump in total bias after batch 17 for the dispersive case.","tokens_in":27400,"tokens_out":1109,"duration_ms":20709,"significance":"If the claims hold, the work supplies a practical, modular tool for tracking slowly drifting multiparameters in high-rate, low-SNR continuous-measurement experiments that are common in superconducting platforms. The explicit derivation of the batch-averaged Kraus map, the pedagogical modular presentation of particle filters versus SMC samplers, and the use of real experimental records (rather than purely synthetic data) are concrete strengths that lower the barrier to adoption. The hyperparameter selection protocol on independent simulations and the open acknowledgment of the Γ–η identification problem further increase the paper’s utility as a methodological reference.","major_comments":[{"comment":"Sections 4.1–4.2 and Eq. (33): The central experimental claim that the SMC estimates are “better” than laboratory calibration rests on lower signal-reconstruction RMSE of the batch-averaged voltage trajectory. Because the observation model (Eqs. 24/29 and the O(Δτ^{2}) map of Appendix A) is incomplete, a systematically biased parameter set can still produce a lower RMSE simply by absorbing residual model error (unmodeled T_{2} dephasing, residual T_{1}, scaling-factor mismatch). The authors themselves document a near-perfect Γ–η degeneracy (Sec. 4.3, Fig. 6). Consequently the RMSE ranking and the reported bias jump cannot be taken as conclusive evidence of multiparameter accuracy without either (i) an independent ground-truth check or (ii) controlled misspecification simulations that quantify how much RMSE improvement can be obtained from compensating bias alone. The language of the abst","section":"Sections 4.1–4.2, Eq. (33), Sec. 4.3, Fig. 6"},{"comment":"Section 2.6 and Appendix A: The batch-averaged Kraus map and the subsequent weight updates assume that the unknown parameters are constant (or negligibly varying) inside each batch of Ns trajectories. The chosen Ns values (10 200 for fluorescence, 4 000 for dispersive) are large enough that any drift on the batch acquisition timescale systematically biases the posterior. The sudden jump recovered after batch 17 in the dispersive data set itself suggests that the constant-within-batch premise can be violated. A quantitative bound on the bias incurred when the premise fails, or an adaptive batch-size procedure, is needed to underwrite the time-dependent claims.","section":"Section 2.6, Appendix A"}],"minor_comments":[{"comment":"Section 3, paragraph after Eq. (19): the sentence “The reasons for the order of operations is threefolds: Firstly, . Secondly, .” is incomplete; the missing clauses should be restored or the sentence deleted.","section":"Section 3"},{"comment":"Figures 2 and 4: the blue error bars and red shaded regions are defined only in the captions; a short legend or explicit statement in the main text would improve readability.","section":"Figures 2, 4"},{"comment":"Notation for the residual offset switches between v_off,t, ˆV_off and V_off without a single clarifying equation; a short glossary or consistent definition would help.","section":"Sections 4.1–4.2"},{"comment":"Appendix B figures: the true-parameter dashed lines are sometimes hard to distinguish from the estimated traces; thicker or differently styled lines would improve clarity.","section":"Appendix B"}],"recommendation":"major_revision","confidential_remarks":"The methodological core (hybrid SMC + efficient batch Kraus map) is solid and publishable. The experimental claims are currently overstated relative to the strength of the validation metric; once the language is tempered and the two load-bearing caveats are addressed, the paper will be a useful contribution. Scope is appropriate for a quantum-information / quantum-control journal."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful core is a hybrid online-offline SMC sampler (SIS with zero-noise transition plus SMC-sampler resampling) combined with an O(Δτ²) batch-averaged Kraus map (Appendix A, Eq. 43). Particle filters and SMC samplers for quantum trajectories already exist (Chase & Geremia, Ralph et al. 2017); batch estimation is classical (Chopin). What is new is the concrete marriage of the two plus the efficient averaged map and the experimental multiparameter trajectories on fluorescence and dispersive-z data.\n\nThey do the practical things well. Hyperparameters are fixed on independent simulations before touching experiment. The modular derivation of SIS and the SMC sampler is clear enough that someone could re-implement. On the fluorescence set the reconstructed average voltage has lower RMSE than the published static calibration (0.37 vs 0.75). On the dispersive set the algorithm recovers a clear jump in total bias after batch 17 that the laboratory calibration never reported. That jump is the strongest empirical claim.\n\nThe soft spots are real but proportionate. The stress-test is right that RMSE of the mean trajectory is not an independent ground-truth test of multiparameter accuracy: the observation model is incomplete (Γ–η product degeneracy is documented in their own Fig. 6; unmodeled T₂ can be absorbed into the same product). A biased parameter set can still match the averaged voltage better than a static calibration. The slow-drift assumption inside each batch is also load-bearing; if parameters move on the batch timescale the posterior is systematically biased. No code or data are released, so reproducibility is only moderate. None of this collapses the paper; it just means the claims should be read as “better reconstruction and previously unseen drift features under the stated model,” not “true multiparameter recovery.”\n\nThis is for people who actually run continuous-measurement experiments on superconducting qubits and need a practical multiparameter tracker. It deserves a serious referee. I would cite the method and the bias-jump observation if I were working in the same niche.","headline":"Solid hybrid SMC + batch-averaged Kraus map that improves reconstruction on two real qubit datasets and surfaces a bias jump calibration missed; RMSE is relative evidence, not ground truth.","tokens_in":27955,"tokens_out":516,"would_cite":true,"duration_ms":5099,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["03.65.Yz","03.65.Wj","42.50.Lc","85.25.Cp"],"model":"grok-4.5","headline":"A hybrid SMC-plus-batch method tracks multiparameter drifts in noisy quantum data better than static calibration, recovering a hidden jump.","keywords":["sequential Monte Carlo","particle filter","batch estimation","continuous quantum measurement","parameter tracking","superconducting qubit","Kraus map","homodyne detection"],"falsifier":"Re-run the identical algorithm on the same fluorescence and dispersive-z data sets with deliberately halved batch size; if the reconstructed RMSE no longer improves over calibration and the bias jump disappears, the batch-constancy premise has failed.","tokens_in":27968,"feed_emoji":"⚛️","tokens_out":776,"duration_ms":7547,"temperature":0.7,"pith_summary":"Real quantum devices suffer parameter drifts, yet their continuous measurement records are too noisy and too fast for ordinary online estimators to keep up. This paper builds a hybrid estimator that first averages successive trajectories into short batches (raising the signal-to-noise ratio) and then feeds the averaged records into a sequential Monte-Carlo particle filter whose resampling step is an SMC sampler. The dynamics inside each batch are evolved with a first-order-averaged Kraus map that the authors derive and approximate for speed. Applied to two published superconducting-qubit datasets, the same algorithm both beats the published calibration numbers on signal-reconstruction error and uncovers a sudden jump in total bias that the original calibration never reported. The practical upshot is a modular, offline-online procedure that can track slow drifts while remaining computationally tractable for the high-rate, low-SNR regime typical of present-day quantum experiments.","feed_headline":"Hybrid SMC tracks qubit drifts better than static calibration","feed_subtitle":"Batch-averaged Kraus maps recover a hidden bias jump and lower reconstruction error on real data","key_machinery":"The batch-averaged Kraus map (exact form Eq. 28, O(Δτ^{2}) approximation Eq. 43) that lets a single particle filter evolve an averaged trajectory while still computing the correct Gaussian likelihood for the mean record.","core_discovery":"When continuous homodyne records from superconducting qubits are partitioned into batches, averaged, and processed by an SMC particle filter whose resampling is itself an SMC sampler, the resulting multiparameter trajectories reconstruct the observed signals more accurately than the published static calibrations and reveal previously undetected jumps in the bias parameter.","pith_inferences":["The same batch-averaged likelihood could be paired with an unscented Kalman filter or variational Bayes recursion, testing whether the performance gain is specific to SMC or generic to any sequential Bayesian update.","Because the method already recovers a jump that calibration missed, it could serve as an online diagnostic for sudden environmental events (flux jumps, TLS flips) in larger quantum processors.","Smoothing the particle trajectories with future data (as the authors themselves flag) would turn the present filter into a smoother and further reduce the residual RMSE."],"forward_implications":["Slow parameter drifts that static calibrations miss become visible in continuous quantum experiments without new hardware.","Signal-reconstruction RMSE supplies a model-independent figure of merit for comparing estimators when true parameter values are unknown.","The modular SMC-plus-batch skeleton can be dropped onto any Markovian continuous-measurement model once an averaged Kraus map is written.","FPGA-friendly simplifications of the same recursion open a path to real-time feedback estimation on nanosecond platforms."],"fun_headline_variants":["Hybrid SMC finds bias jumps missed by static qubit calibration","Batch-averaged SMC tracks multiparameter drifts better than calibs","Online-offline SMC recovers hidden jumps on superconducting qubits","SMC with batch Kraus maps beats static calibration on real signals","Hybrid Monte Carlo uncovers parameter jumps static methods miss"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"Parameters are treated as constant inside each batch so that a single averaged Kraus map remains valid; any drift on the batch timescale systematically biases the posterior.","fun_headline_variants_meta":{"raw":{"variants":["Hybrid SMC finds bias jumps missed by static qubit calibration","Batch-averaged SMC tracks multiparameter drifts better than calibs","Online-offline SMC recovers hidden jumps on superconducting qubits","SMC with batch Kraus maps beats static calibration on real signals","Hybrid Monte Carlo uncovers parameter jumps static methods miss"]},"model":"grok-4.5","effort":"low","cost_usd":0.00366,"raw_usage":{"total_tokens":1182,"prompt_tokens":762,"num_sources_used":0,"completion_tokens":82,"cost_in_usd_ticks":36600000,"prompt_tokens_details":{"text_tokens":762,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":338,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":762,"tokens_out":82,"duration_ms":3254,"temperature":1.0,"reasoning_tokens":338,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T05:36:31.287775+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the identical algorithm on the same fluorescence and dispersive-z data sets with deliberately halved batch size; if the reconstructed RMSE no longer improves over calibration and the bias jump disappears, the batch-constancy premise has failed.","supporting_citations":[],"review_version":1}