{"id":"f625bd4b-0311-42b1-bf8d-28ba8e7a859a","arxiv_id":"2412.16092","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":13,"one_line_summary":"A hybrid Lindblad-master-equation noise model with 10 parameters per qubit and 3 per pair predicts RB, dynamical-decoupling, and H2 VQE dynamics on IBM transmon hardware, reaching 0.5% relative energy error at the optimal bond length.","lead":"This paper builds a sparse noise model for IBM quantum computers that tracks slow, correlated noise sources as well as simple ones. The model predicts real hardware behavior on tests like a hydrogen-molecule energy calculation, beating IBM's default error model by about seven times at one key point.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"VQE headline (0.5% error, 7x improvement) rests on a single unrepeated data point with no error bars and a stale default-model comparison; the paper's own stability data show parameter drift, so the central quantitative claim is not yet robust.","rationale":"The paper makes a plausible and internally detailed argument: a sparse LME extended with TLS, crosstalk, and low-frequency stochastic noise is fit to seven characterization experiments and tested on RB, CPMG, DD, and VQE. The analytical derivations in the appendices are substantial, and the out-of-sample RB and CPMG predictions provide real evidence for the model's usefulness. I considered whether the FTTPS PSD reconstruction loop (fitting and then predicting the same FTTPS data) is the weakest point, but the later CPMG and DD experiments provide independent validation of the reconstructed spectra, so that circularity is not fatal. I also considered model completeness: the assumed families in Eqs. (1)-(10) could miss drive-induced TLS dynamics or non-ZZ crosstalk, but the paper's demonstrations on several qubits and two-qubit gates make this a scaling concern rather than a demonstrated failure at the tested scale. The single most load-bearing issue is the quantitative headline itself: a one-point, no-error-bar comparison whose baseline is a stale calibration model. The authors' own stability data (Appendix I) show that parameters drift on the timescale of the experiments, so the claimed 0.5% precision is not shown to be reproducible. A repeated, properly controlled VQE campaign with a freshly characterized default model would settle whether the 0.5% and 7x figures are robust or artifacts of a single favorable run. This supports keeping the reader's CONDITIONAL verdict rather than upgrading it.","tokens_in":48031,"tokens_out":4736,"duration_ms":49234,"concrete_test":"Repeat the VQE H2 experiment on the same device (e.g., ibm_algiers qubits 12 and 15) in at least 3 independent campaigns spread over ≥24 hours, running the full characterization protocol immediately before each VQE measurement and refreshing the IBM backend properties at the same time. For each campaign, compute the relative error Δ(R) from Eq. (25) at least 5 bond lengths and report the mean and standard error; also recompute predictions using parameters perturbed by the observed drift in Appendix I. If the 0.5% error is not reproduced in the majority of campaigns and the improvement over a freshly characterized IBM default model is not consistently ≥5x, the headline quantitative claim should be weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest quantitative claim — that the model predicts VQE expectation values within 0.5% relative error, a 7x improvement over default hardware noise models (Sec. V B, Fig. 9) — is supported by one experimental point at one bond length (Ropt = 0.75 Å), with no repeated runs, no error bars, and no sensitivity analysis. The comparison is also confounded by calibration age: the default IBM model used backend properties collected roughly 9 hours before the experiment, whereas the proposed LME model was freshly characterized. The authors do include a 'Markovianized' version of their model to isolate non-Markovian contributions, but that also yields only a single point (Δ ≈ 3.8%) without uncertainty. Appendix I (Fig. 17) shows that fitted parameters drift within a one-hour window on a single qubit — e.g., TLS coupling ξ varies by roughly 0.2 MHz and q, β, ν show visible trends — so the assumption that parameters learned at characterization time transfer to the later VQE run is not established at the claimed precision. The central claim that the model is predictive at the 0.5% level therefore rests on a favorable snapshot rather than a demonstrated reproducible property.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a sparse, hybrid noise model for IBMQ transmon devices that combines Lindblad master equation terms (local GAD, phase damping, bit-flip control noise) with extended-Markovian degrees of freedom (TLSs, spectator crosstalk) and classical stochastic noise (dephasing and control PSDs). The model is parameterized by about ten parameters per qubit and three per qubit pair, learned from seven characterization experiments (SPAM, T1, T2, Ramsey, FTTPS, FPW, XT, CR). The authors validate the model on multiple IBMQ devices, showing agreement for Markovian qubits, TLS-induced Ramsey beats, crosstalk, correlated dephasing and control noise, and ECR gates. They then use the model to predict randomized benchmarking decay, multi-qubit dynamical decoupling curves, and a VQE dissociation curve for H2, claiming a 0.5% relative error in energy, a 7x improvement over the default IBM noise model.","tokens_in":48422,"tokens_out":3001,"duration_ms":27248,"significance":"If the results hold, the paper makes a useful contribution: it demonstrates that a physically motivated, low-parameter noise model can predict out-of-sample hardware dynamics across several experiment classes. The analytical Bloch-vector solutions in the appendices are checked against LME simulations (Fig. 10), the RB prediction in Fig. 2(e) is a genuine out-of-sample test, and the CPMG predictions in Figs. 5(c,d) use PSDs extracted from FTTPS, a different experiment class, which is a strong validation design. The multi-qubit DD and VQE demonstrations extend the model beyond single-qubit benchmarking. The reproducibility of the characterization protocol and the explicit parameter sparsity are also strengths. However, the headline quantitative claim (0.5% VQE error) is supported by a single unrepeated data point, and the paper's own stability data indicate parameter drift, so the central predictive claim is not yet robust as stated.","major_comments":[{"comment":"The headline claim of 0.5% relative error and 7x improvement over the default IBM model rests on a single experimental point at Ropt = 0.75 Å, with no repeated runs, no error bars, and no sensitivity analysis. In addition, the comparison is confounded by calibration age: the default IBM model uses backend properties collected roughly 9 hours before the experiment, while the proposed LME model was freshly characterized. Appendix I (Fig. 17) shows that fitted parameters drift within a one-hour window on a single qubit (e.g., TLS coupling xi varies by ~0.2 MHz; q, beta, and nu show visible trends). Thus, the assumption that characterization-time parameters transfer to the later VQE run is not established at the claimed 0.5% precision. I request either repeated VQE runs with statistics and error bars, a time-matched default model comparison, or a quantitative sensitivity analysis showing that parameter drift does not materially change the reported error.","section":"Sec. V B, Fig. 9"},{"comment":"The PSD reconstruction for qubit 4 of ibm hanoi is obtained from FTTPS data and then used to 'predict' the same FTTPS data in the main panel of Fig. 5(a). This is an internal consistency check rather than an out-of-sample validation. The genuinely out-of-sample validation is the CPMG prediction in Figs. 5(c,d), which uses PSDs from FTTPS on different circuits. The text should clearly separate these two roles and avoid implying that Fig. 5(a) provides predictive evidence; currently the narrative could mislead a reader into thinking the FTTPS prediction is independent.","section":"Sec. IV C, Fig. 5(a)"},{"comment":"The model assumes every TLS is initialized in the |+> state at the start of each experiment and couples only via static ZZ interactions. This is a load-bearing and ad hoc assumption: any other initial TLS state would change the Ramsey and DD predictions. The paper provides no direct experimental test of this initial condition, and Appendix H shows that the distinction between TLS and crosstalk/detuning frequencies is resolved using a specific modeling assumption (beta = -J). I recommend either a direct test (e.g., varying TLS preparation or checking consistency across different experiments) or a sensitivity analysis showing that plausible deviations from |+> initialization do not affect the main predictions.","section":"Sec. II C, Eq. (8) and Appendix H"}],"minor_comments":[{"comment":"There is a typographical error: 'Krauss operators' should be 'Kraus operators'.","section":"Appendix D"},{"comment":"The text says 'For qubit 0 [see Fig. 5(d)]', but Fig. 5(d) corresponds to qubit 2; the reference should likely be to Fig. 5(c) or to qubit 2.","section":"Sec. IV C 1, Fig. 5(c,d)"},{"comment":"The relative error inset of Fig. 9 would benefit from error bars or shaded confidence intervals, especially since the experimental energies are estimated from 10,000 shots. The current single-point claim is difficult to assess without uncertainty quantification.","section":"Sec. V B and Fig. 9"},{"comment":"The statement 'although more experimental data may be needed in order to refine the model' is a limitation that should be made more prominent in the main text, as it qualifies the ECR characterization results.","section":"Sec. IV D"},{"comment":"The notation in Eq. (A21) for the k=0 case ('cos(2 beta tau)') seems inconsistent with the earlier Ramsey expression in Eq. (15), where the argument is (beta tau) not (2 beta tau). Please clarify the prefactors.","section":"Appendix A 4 e, Eq. (A21)"},{"comment":"The definition of the PSD as S_f(omega) = integral_0^tau C_f(t) e^{-i omega t} dt uses a finite-time window, while later uses infinite limits; the windowing conventions should be stated consistently, perhaps with a note about the large-tau limit.","section":"Sec. II D"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the scope of the journal and presents a genuinely useful modeling framework with several strong out-of-sample validations. The main risk is the VQE headline claim: it is supported by a single unrepeated point and the calibration-age confound, and the paper's own stability data undermine confidence in parameter transfer. I would like to see the VQE claim either strengthened with repeated runs/time-matched comparison/error bars or downgraded in the abstract and conclusion to reflect its preliminary nature. The TLS initial-state assumption is a second point that deserves a robustness check. My recommendation is major revision rather than rejection, because the central modeling framework appears sound and the load-bearing issue is fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nHere is my read on arXiv:2412.16092. The paper proposes a sparse Lindblad noise model for transmon qubits, extended with TLSs, crosstalk, and stochastic dephasing/control noise. The main thing to know: the model itself is a sensible integration of known ingredients, and the out-of-sample validation is mostly genuine, but the headline VQE claim (0.5% relative error, 7x improvement) is the softest part of the paper and should not be taken at face value.\n\nWhat is actually new is the specific 10+3 parameter model, the R-FTTPS sequence for isolating correlated control noise, and the systematic characterization across 39 qubits on seven IBM devices. The paper is unusually thorough analytically: the Bloch-vector solutions for all the characterization circuits are derived in the appendices and validated against numerical simulation. The RB prediction is a true out-of-sample test, and the CPMG predictions use PSDs reconstructed from a different experiment class (FTTPS), which is the right experimental logic. The multi-qubit DD results are also genuinely predictive.\n\nThe soft spots. The VQE result is a single experimental point at one bond length, with no repeated runs, no error bars, and no sensitivity analysis. The comparison to the IBM default model is also confounded by calibration age: the default model used properties collected roughly 9 hours before the experiment, while the proposed LME model was freshly characterized. The paper's own stability data (Appendix I) shows fitted parameters drifting over one hour, so the 0.5% number is a favorable snapshot rather than a demonstrated reproducible property. The FTTPS PSD reconstruction loop in Fig. 5(a) is partly circular, but the subsequent CPMG prediction is out-of-sample, so this is minor. The model's completeness (no unmodeled context-dependent noise) is a load-bearing assumption, but that is the standard caveat for effective noise models.\n\nOverall the central argument holds up: this is a useful engineering-oriented framework for predictive noise modeling. The VQE claim needs more data and a fairer baseline before it can be the paper's centerpiece.\n\nThis paper deserves a serious referee. The referee should focus on the VQE section, ask for repeated runs and error bars, and for a sensitivity analysis on the learned parameters. I would not desk-reject it.\n\nBest,\n\n[You]","headline":"Solid effective noise modeling paper with a strong out-of-sample core and an over-claimed VQE headline that needs error bars and repetition.","tokens_in":48998,"tokens_out":3006,"would_cite":true,"duration_ms":27238,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A sparse noise model built from a Lindblad master equation with only ten parameters per qubit predicts IBM transmon hardware dynamics to within 0.5% relative error, a sevenfold improvement over the platform's default noise model.","keywords":["non-Markovian noise","noise model","transmon qubits","Lindblad master equation","quantum noise spectroscopy","cross-resonance gate","two-level systems","variational quantum eigensolver"],"falsifier":"Run the H2 VQE circuit on a qubit pair whose FTTPS shows a high-frequency resonance peak (as in the paper's Appendix G) and check whether the model still predicts the experimental expectation values within 0.5% relative error; a failure would show the static-ZZ, single-TLS ansatz is incomplete. A second direct test: calibrate on one qubit pair and apply the model to a different pair on the same device; if the 0.5% accuracy does not hold, the parameters are not context-stable.","tokens_in":47783,"feed_emoji":"⚛️","tokens_out":8841,"duration_ms":72002,"temperature":0.7,"pith_summary":"The paper claims that a physically motivated noise model with only ten parameters per qubit and three per qubit pair can capture and predict both Markovian and non-Markovian noise in IBM transmon devices. The model is built from a Lindblad master equation extended with classical stochastic dephasing and control noise, plus quantum degrees of freedom for two-level systems and spectator-qubit crosstalk. From seven short characterization experiments, the learned parameters predict randomized benchmarking error rates, state-dependent dynamical decoupling decays, and a variational quantum eigensolver dissociation curve for H2. As a hardware proxy the model predicts expectation values within a relative error of 0.5%, a sevenfold improvement over the platform's default noise model. If correct, this offers a practical route to hardware-accurate circuit simulation without expensive tomographic characterization.","feed_headline":"10-parameter noise model predicts IBM qubit behavior to 0.5%","feed_subtitle":"Learned from seven short experiments, the model beats the default hardware error model sevenfold in a VQE hydrogen simulation.","key_machinery":"The central object is the modified Lindblad master equation $$\\dot{\\rho}(t)=-i[H(t),\\rho]+\\sum_k \\gamma_k (L_k \\rho L_k^\\dagger - \\frac{1}{2}\\{L_k^\\dagger L_k,\\rho\\})$$ with Hamiltonian $H = H_C + H_N + H_{XT} + H_{TLS}$, where $H_C$ is single- and two-qubit control, $H_N(t)=\\sum_j [\\epsilon_j(t)H_C^{(j)}(t)+\\beta_j(t)\\sigma_z^{(j)}/2]$ is Gaussian wide-sense-stationary stochastic dephasing and control noise, $H_{XT}$ is ZZ crosstalk to spectator qubits, and $H_{TLS}$ couples each data qubit to one or more two-level systems initialized in the $|+\\rangle$ state via a static ZZ coupling. The dissipator contains generalized amplitude damping, phase damping, and a bit-flip channel active during x-rotations. The argument is carried by (i) analytical Bloch-vector solutions of the LME for each characterization circuit, (ii) the filter-function formalism with fixed-total-time pulse sequences (FTTPS) to reconstruct the dephasing and control power spectral densities $S_\\beta(\\omega)$ and $S_\\epsilon(\\omega)$, and robust FTTPS (R-FTTPS) to isolate control noise, and (iii) a channel-reduced operator-sum representation for scalability.","core_discovery":"The authors establish that a hybrid noise model—local Markovian dissipation (generalized amplitude damping, phase damping, and a control bit-flip channel) plus extended Markovian degrees of freedom (ZZ crosstalk to spectator qubits and ZZ-coupled two-level systems initialized in the |+> state) plus Gaussian wide-sense-stationary stochastic dephasing and control noise—captures and predicts a wide range of single- and two-qubit behaviors on IBM transmon devices. The model's parameters are learned from seven noise-amplification experiments (T1, T2 Hahn echo, Ramsey, SPAM, fixed-total-time pulse sequences, finite-pulse-width sequences, crosstalk, and cross-resonance gates). The learned model predicts error-per-Clifford rates in randomized benchmarking, state-dependent decays under multi-qubit XY4 dynamical decoupling, and the H2 dissociation curve from a variational quantum eigensolver. The headline quantitative claim is that, as a training proxy for hardware, the model predicts expectation values within a relative error of 0.5%, a 7x improvement over the platform's default noise model; replacing the correlated-dephasing term with a Markovian phase-damping channel degrades the accuracy to about 3.8%, showing the non-Markovian terms carry the improvement.","pith_inferences":["If the predictive accuracy generalizes to larger systems, it suggests a physics-informed sparse ansatz can outperform generic dense noise models in the few-qubit regime, because the dominant error mechanisms in fixed-frequency transmons are already captured by this small parameter set.","The framework—Markovian dissipators plus classical stochastic noise plus a few quantum TLS/spectator degrees of freedom—should transfer to other qubit modalities where dephasing and TLS coupling dominate, though the paper only demonstrates it on IBM fixed-frequency transmons.","The Gaussian wide-sense-stationary assumption could be tested by measuring higher-order noise cumulants (e.g., via spin-locking or higher-order QNS); if TLS telegraph switching contributes appreciably, the model would need non-Gaussian extensions.","A practical test of the method's value: calibrate the model on one device generation and apply it to another; if characterization transfer fails, the model's usefulness as a universal hardware proxy is limited."],"forward_implications":["For the 64% of qubits that are purely Markovian, the model provides a full error description from a handful of short experiments, predicting randomized benchmarking error rates without running RB.","For qubits with correlated dephasing (26%), the model predicts how much fidelity improves with dynamical decoupling pulse number, identifying which qubits will benefit most from DD.","The VQE result implies variational algorithms can be trained offline against a faithful noise model, saving quantum hardware time and enabling noise-aware ansatz selection before submitting circuits.","The operator-sum channel reduction extends the model beyond a few qubits, so hardware-accurate simulation may scale to larger devices by composing per-qubit and per-pair channels.","Two-qubit ECR gates require no two-qubit dissipative terms; single-qubit dissipation plus coherent Hamiltonian corrections suffice, simplifying future multi-qubit characterization."],"supporting_citations":[{"why":"Supplies the ARMA-based FTTPS quantum noise spectroscopy method used to reconstruct dephasing and control power spectral densities from fixed-total-time pulse sequences.","marker":"[58]"},{"why":"Provides the quantum noise spectroscopy framework for characterizing spatio-temporally correlated noise, the basis for the correlated dephasing analysis.","marker":"[57]"},{"why":"Formalizes the filter-function formalism that links noise power spectral densities to decay rates, used throughout the stochastic noise model and channel reduction.","marker":"[93]"},{"why":"Introduces the approach of coupling qubits to additional quantum degrees of freedom in a Lindblad master equation, which the paper extends to TLS modeling.","marker":"[72]"},{"why":"Supplies the echo cross-resonance gate error model and effective Hamiltonian that the two-qubit noise model uses.","marker":"[102]"},{"why":"Defines the ECR gate as implemented on IBM hardware, the two-qubit operation characterized in the paper.","marker":"[104]"},{"why":"Provides randomized benchmarking, the validation protocol used to test the model's predictive power against hardware.","marker":"[95]"},{"why":"Supplies the H2 operator-to-circuit mapping used in the variational quantum eigensolver demonstration.","marker":"[125]"},{"why":"Models non-Markovian noise in driven superconducting qubits, supporting the TLS interpretation of multi-frequency Ramsey oscillations.","marker":"[112]"}],"fun_headline_variants":["10-parameter model predicts IBM qubits to 0.5%","Sparse non-Markovian model beats IBM's default error model 7x","7 experiments train a 10-parameter model that predicts IBM quantum hardware","Non-Markovian noise model predicts IBM transmon dynamics to 0.5%","Sparse model from 7 experiments predicts IBM qubit behavior 7x better than default"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model assumes every relevant noise process on the studied devices belongs to one parameterized family—local damping and bit-flip, static detuning, ZZ crosstalk, two-level systems with static coupling, and Gaussian wide-sense-stationary dephasing and control noise—and that parameters learned from short calibration circuits transfer to arbitrary circuits for the duration of an experiment.","fun_headline_variants_meta":{"raw":{"variants":["10-parameter model predicts IBM qubits to 0.5%","Sparse non-Markovian model beats IBM's default error model 7x","7 experiments train a 10-parameter model that predicts IBM quantum hardware","Non-Markovian noise model predicts IBM transmon dynamics to 0.5%","Sparse model from 7 experiments predicts IBM qubit behavior 7x better than default"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001264,"raw_usage":{"total_tokens":5219,"prompt_tokens":1031,"completion_tokens":4188,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":647,"completion_tokens_details":{"reasoning_tokens":4081}},"tokens_in":647,"tokens_out":4188,"duration_ms":27970,"temperature":1.0,"reasoning_tokens":4081,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:47:47.296793+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the H2 VQE circuit on a qubit pair whose FTTPS shows a high-frequency resonance peak (as in the paper's Appendix G) and check whether the model still predicts the experimental expectation values within 0.5% relative error; a failure would show the static-ZZ, single-TLS ansatz is incomplete. A second direct test: calibrate on one qubit pair and apply the model to a different pair on the same device; if the 0.5% accuracy does not hold, the parameters are not context-stable.","supporting_citations":[{"cited_title":"Knill, D","cited_arxiv_id":null,"evidence_quote":"Defines the ECR gate as implemented on IBM hardware, the two-qubit operation characterized in the paper."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides randomized benchmarking, the validation protocol used to test the model's predictive power against hardware."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Models non-Markovian noise in driven superconducting qubits, supporting the TLS interpretation of multi-frequency Ramsey oscillations."}],"review_version":1}