Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Enhanced Qubit Readout via Reinforcement Learning

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A reinforcement-learning agent discovers readout waveforms that complete qubit measurement and resonator reset in under 700 ns, about three times faster than the default configuration, with assignment error as low as 4.6e-3.

desk verdict A genuinely useful hardware result—RL-cut readout/reset times by 2-3x on real IBM devices—but the fidelity model in Eq. (6) is not a valid probability model and the 'state-of-the-art' fidelity claim overreaches. read the letter →

arxiv 2412.04053 v3 pith:7FM2ORB2 submitted 2024-12-05 quant-ph

classification quant-ph
keywords reinforcementlearningqubitreadoutsuperconductingqubitsdispersiveresonatorresetPPOwaveformoptimizationactivefour-tone
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Readout is often the slowest and noisiest step in superconducting quantum computation, and this paper asks whether a machine-learning agent can design the entire readout waveform rather than tuning individual segments. The authors train a model-free reinforcement-learning agent in a quasiclassical Langevin simulator whose parameters are measured from two cloud quantum processors, then transfer the learned pulse to real hardware. They report assignment errors as low as $(4.6 \pm 0.4)\times10^{-3}$ with total readout-plus-reset times of 470 ns and 675 ns on the two test devices, roughly three times faster than the devices' default 1400 ns configuration at comparable fidelity. The learned waveforms are stable under $\pm 10\%$ drifts in $\kappa$ and $\chi$ and collapse onto a simple analytical four-tone form, A4R, that can be calibrated with standard measurements. If these transfers hold, the result matters because faster, high-fidelity readout directly accelerates error correction and mid-circuit measurement, which otherwise idle a quantum processor.

What carries the argument

The load-bearing machinery is the PPO agent interacting with a quasiclassical Langevin environment. The environment evolves two coherent amplitudes $\alpha_g(t)$ and $\alpha_e(t)$ under $\dot{\alpha}_{g/e}(t) = -(\kappa/2 \mp i\chi)\alpha_{g/e}(t) - iA(t)$, converts their separation into a time-dependent assignment fidelity through a Gaussian-noise model with a fitted scale, initialization fidelity, and photon-number-dependent qubit decay rates, and returns a reward that penalizes slow reset, rough or nonzero terminal amplitudes, and photon populations above the default readout's photon number. The discovered structure, A4R, is the named central output: four segments, namely ring-up, steady-state readout, photon depletion, and kickback, each with an amplitude and a duration, with the ring-up and depletion amplitudes set to the hardware maximum and their durations given by closed-form expressions in $\kappa$.

What would settle it

Run the optimized pulse on a third device whose $\kappa/\chi$ lies outside the two tested values, or intentionally detune $\kappa$ or $\chi$ by more than 10%, and compare the measured single-shot assignment error and the resonator photon population after 470-675 ns against the paper's simulated predictions; a mismatch beyond shot noise would indicate that the surrogate environment, not the reinforcement-learning procedure, is the source of the reported speedup. A second check is to replace Eq. (1) with a full master-equation simulation that includes measurement-induced dephasing and ionization, retrain the agent, and see whether the discovered waveform and its assignment error change.

Watch

Extended reading notes

Core claim

The central claim is that a reward function combining assignment fidelity, total reset time, pulse smoothness, and a photon-number cap is enough for a reinforcement-learning agent to rediscover a near-optimal readout protocol: a high-amplitude ring-up that drives the resonator toward steady state quickly, a short calibrated readout segment that reaches maximum separation before steady state, and a two-tone active reset that empties the resonator to below 0.05 photons. The paper argues that the optimal measurement time is not the steady-state time, because qubit decay and measurement-induced transitions accumulate during longer pulses, so the best fidelity occurs at an intermediate time, namely 264 ns on one device. It further claims that the learned waveform can be compressed into eight parameters, four segment amplitudes and four durations, and that this A4R pulse matches the reinforcement-learning pulse's assignment error and speed on hardware while remaining stable under realistic parameter drifts.

Load-bearing premise

The whole optimization rests on the fitted quasiclassical surrogate, which uses two independent coherent trajectories, a Gaussian resonator state, and a constant-noise fidelity formula calibrated on the same two devices, transferring faithfully to real hardware; if that surrogate misses measurement-induced effects outside the tested parameter range, the claimed speed and stability may not generalize.

Editorial extensions

If this is right

  • If the central claim is correct, readout need not wait for resonator steady state: reaching maximum signal-to-noise at an intermediate time shortens the measurement without sacrificing assignment fidelity.
  • Active two-tone reset can clear tens of photons to below 0.05 in hundreds of nanoseconds, so repeated measurements can be chained with far less idle dead time.
  • The A4R form means the optimized behavior can be deployed without a neural network on any comparable device, using only a few calibration scans.
  • The demonstrated stability under $\pm 10\%$ drifts in $\kappa$ and $\chi$ implies the pulse does not need continuous re-optimization over a typical device-drift timescale.
  • The same reward structure should apply to other dispersive-readout hardware, because the training environment only needs $\kappa$, $\chi$, photon number, and decay rates.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the same training loop could be pointed at nonlinear readout regimes, such as large self-Kerr or near-ionization photon numbers, where no analytical four-tone ansatz exists, turning reinforcement learning from a rediscovery tool into a discovery tool.
  • Because the reset durations in A4R are essentially set by $\kappa$, the protocol could be ported to a new device with only a $\kappa$ measurement, even when $\chi$ is poorly known.
  • A natural test would benchmark A4R inside a dynamic quantum circuit with many mid-circuit measurements, predicting overall runtime savings that scale roughly with the number of measurements.
  • The surrogate's constant-noise, Gaussian-state fidelity formula has not been validated outside the tested parameter window, so the robustness claim should not be extrapolated to much higher photon numbers or vastly different $\kappa/\chi$ without revalidation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a reinforcement learning (RL) framework, based on PPO, to optimize the full dispersive readout pulse for superconducting qubits. The agent is trained against a quasiclassical Langevin simulator whose parameters are calibrated to IBM devices, and the resulting waveforms are executed on IBM Kyoto and Brisbane. The authors report assignment errors of (4.6 ± 0.4) × 10^-3 on Kyoto and (7.0 ± 1.0) × 10^-3 on Brisbane, with total readout-and-reset durations of 470 ns and 675 ns, respectively, compared with 1400 ns for the default configuration. They also identify a four-segment analytical form called A4R and claim robustness against ±10% parameter drifts, supported by simulation and by a comparison with the CLEAR protocol.

Significance. If the central speedup claim holds, the work is practically valuable: reducing the measurement-and-reset time while preserving assignment fidelity is directly relevant for quantum error correction and mid-circuit measurement. The paper's strengths are the live hardware measurements with error bars on two IBM devices, the openly stated code and parameters, the short training time, and the reduction of the learned waveform to a simple calibration-ready analytical form. However, the fidelity improvement over the default protocol is not statistically significant in the reported data, and the simulated robustness and CLEAR-comparison results rely on a fidelity model that is not a valid probability model. The claimed generality and state-of-the-art fidelity therefore need additional support before the paper can be accepted.

major comments (3)
  1. [Appendix B, Eq. (6)] The fidelity model F(t) = 0.5[1 + erf(F0 × λS(t) × Fq(t))] is not a valid assignment-fidelity probability. Since the erf argument grows without bound as S(t) → ∞, F approaches 1 even for a fixed, non-unit survival factor Fq; physically, a fraction 1 - Fq of excited-state preparations that decay into the ground-state channel should impose a high-SNR assignment error of at least (1 - Fq)/2, independent of S. The model also gives F(0) = 0.5 for any initialization fidelity, and F0 and λ appear only as a product, so the Appendix B fit is underdetermined. Because the reward in Eq. (3) maximizes max_t F(t), the PPO agent may optimize against this artifact rather than against the true measurement physics. The live hardware results at tFmax = 264 ns (Kyoto) and 542 ns (Brisbane) partly mitigate the concern, as Fq is close to unity there, but the simulated robustness landscapes in Fig. 3(d,e) and the CLEAR comparison in Table IV are generated using Eq. (6) and therefore inherit the error. I request a corrected fidelity model—for example, an explicit mixture of ground and excited readout distributions with decay—or a validation of the surrogate against a full master-equation treatment, and a rerun of the affected simulation-based claims with that corrected model.
  2. [Section IV.B and Table II] The claim that the RL waveform achieves 'equal or slightly superior fidelities' is not supported by the quoted uncertainties. On Kyoto, (4.6 ± 0.4) × 10^-3 versus (5.8 ± 0.9) × 10^-3 is within roughly 1.2 combined standard deviations, and on Brisbane the RL value (7.0 ± 1.0) × 10^-3 is identical to the default value to the reported precision. The data therefore establish only that the RL pulse attains comparable fidelity while being substantially faster. The abstract and Section IV.B should be revised to state 'comparable fidelity' rather than 'state-of-the-art performance' or 'slightly superior fidelities,' unless a significance test is provided.
  3. [Section III and Appendix B (training environment)] The training environment is a quasiclassical surrogate whose free parameters—λ, F0, γ0, γP, the reward coefficients ki, and the reset penalty factor m—are fitted to the same devices that are used for hardware validation. This is not circular in the sense of using the final hardware result to define the objective, but it creates a risk that the policy is optimized to idiosyncrasies of the self-calibrated model rather than to robust measurement physics. The paper would be strengthened by an explicit transfer experiment to a device with parameters outside the fitted range, or by an analysis of the sensitivity of the measured performance to the fitted parameters. As written, the assertion that the method is 'readily applicable to generic superconducting devices' rests on only two test points and on a surrogate whose validity is questioned by the issues in Eq. (6).
minor comments (5)
  1. [Abstract and Section IV.B] The phrase 'almost three times faster' is accurate for Kyoto (1400 ns to 470 ns) but for Brisbane the speedup is about 2.1×; please specify per-device factors or state 'two- to threefold.'
  2. [Fig. 3(a) caption] The labels 'RL wfA4R Square' and 'RL wfA4R Square' are ambiguous; please use distinct labels for the learned waveform, the analytical A4R pulse, and the default square pulse, and apply the same labels consistently across all panels.
  3. [Appendix D, after Eq. (8)] The derivation of τ1 is described only as using 'first-order approximations in χt'; please write out the explicit expression that leads to Eq. (8) so the reader can verify the approximation.
  4. [Table IV caption] The column heading 'RL/A4R' is undefined; state that these are simulated values for the RL/A4R waveform, and clarify whether the same fidelity metric and integration weights are used as in the hardware measurements.
  5. [GitHub reproducibility] Reference [24] states that source code and parameters are provided, but the availability of the exact trained network weights and the calibration script for A4R is not explicitly confirmed; please state what is included in the repository.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: headline figures are live hardware measurements; the fitted surrogate model is an input to training but not the basis of the central claims.

full rationale

The central results—assignment error (4.6±0.4)×10^-3, reset times 470/675 ns, and the comparison with default and CLEAR readout—are obtained from execution on IBM Kyoto/Brisbane and from repeated hardware fidelity measurements, not from the training model. The RL reward in Eq. (3) maximizes F(t) from Eq. (6), whose parameters (F0, λ, γ0, γP) are fitted to device calibration data; this is a surrogate used for training, and the paper independently verifies the learned waveform on hardware, so no "prediction" reduces by construction to its fitted inputs. The A4R analytical form is presented as a post hoc distillation of the converged RL waveform (Appendix D), not imposed as an ansatz before training. Robustness landscapes and Table IV are simulations using the same Langevin model, but they are clearly labeled as simulations rather than hardware predictions; any model-validity concerns about Eq. (6) are correctness risks, not circularity. Citations are to standard external results (Blais et al., Gambetta et al., CLEAR); there is no load-bearing self-citation chain or uniqueness argument. The paper benchmarks against external baselines (default IBM pulse and CLEAR) and supplies code for reproducibility.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

No invented physical entities are introduced. The free parameters are calibration or fit constants and reward weights, while the central physical assumptions are the quasiclassical readout model and its transfer to hardware.

free parameters (6)
  • SNR scaling constant lambda = not stated
    Scales the phase-space separation S(t) into fidelity in Eq. 4; fitted to measured assignment infidelities versus acquisition time (Appendix B).
  • qubit initialization fidelity F0 = not stated
    Multiplies FSNR in Eq. 6; extracted by fitting Eq. 6 to device data (Appendix B).
  • qubit decay rates gamma0 and gammaP = not stated
    Set qubit survival Fq(t) in Eq. 5 with photon-population-dependent decay; fitted to measured infidelities (Appendix B).
  • reward coefficients k1-k6 = k1=10, k2=2, k3=1, k4=100, k5=100; k6 not specified
    Chosen by hand and grid sweep to prioritize fidelity over speed and smoothness (Appendix B).
  • reset penalty factor m = 8
    T1 scaling factor in Eq. 7, chosen so the agent actively resets rather than waiting for passive decay (Appendix B).
  • A4R segment parameters A1-A4 and tau1-tau4 = device-specific, not tabulated
    Eight parameters defining the generalized pulse; calibrated per device through sweeps and formulas in Appendix D.
assumptions (5)
  • domain assumption Dispersive readout is governed by the coherent Langevin equation (Eq. 1) with independent amplitudes for |g> and |e> and no nonlinear or ionization dynamics.
    Used throughout training; Appendix B states real pulses, low self-Kerr, and N << Nc, but this excludes measurement-induced state transitions beyond a linear decay term.
  • domain assumption The resonator state can be represented as a Gaussian with constant amplifier noise, so fidelity follows FSNR(t) = 0.5(1 + erf(lambda S(t))).
    Appendix B, Eq. 4; the noise floor is not measured shot-by-shot, it is absorbed into the fitted lambda.
  • domain assumption Qubit survival during readout decays as exp(-gamma0 t - gammaP integral N dtau) with rates fitted to the same devices.
    Appendix B, Eq. 5; used to set the optimal measurement duration in the reward.
  • domain assumption The default IBM square pulse with 680 ns passive delay is an appropriate baseline for the speed comparison.
    Table II; the threefold speedup is measured against this default configuration, not against all possible readout protocols.
  • domain assumption A waveform optimized in the fitted quasiclassical simulator transfers to real hardware without additional closed-loop adjustment.
    The paper tests transfer on two qubits, but the agent itself never sees live hardware; this transfer step is load-bearing.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Enhanced Qubit Readout via Reinforcement Learning." pith.science (2026). https://pith.science/paper/7FM2ORB2

@misc{pith2026241204053,
  author       = {Pith},
  title        = {Pith review of: Enhanced Qubit Readout via Reinforcement Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7FM2ORB2}},
  note         = {Machine review of arXiv:2412.04053}
}
abstract

Measurement is an essential component of robust and practical quantum computation. For superconducting qubits, the measurement process involves the effective manipulation of the joint qubit-resonator dynamics, and it should ideally provide the highest quality for qubit state discrimination with the shortest readout pulse and resonator reset time. Here, we harness model-free reinforcement learning (RL), together with a tailored training environment, to achieve this multi-pronged optimization task. Using the IBM quantum device, we demonstrate that the pulse obtained by the RL agent not only successfully achieves state-of-the-art performance, with an assignment error of $(4.6 \pm 0.4)\times10^{-3}$, but also executes the readout and the subsequent resonator reset almost three times faster than the system's default process. Furthermore, the learned waveforms are robust against realistic parameter drifts and follow a generalized analytical form, making them readily implementable in practice with no significant computation overhead. Our results provide an effective readout strategy to boost the performance of superconducting quantum processors and demonstrate the prowess of RL in providing optimal and experimentally informed solutions for complex quantum information processing tasks.

Figures

Figures reproduced from arXiv: 2412.04053 by the authors.

Figure 1
Figure 1. FIG. 1. Depiction of the learning process of a proximal policy optimization (PPO) agent interacting with our environment. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 1
Figure 1. Compared to other numerical algorithm–such as [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2. RL learning curves and resulting waveforms [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figures from the paper (2 more)
Figure 3
Figure 3. Figure 3: FIG. 3. Performance and stability of optimized waveforms discovered via RL. (a) An exemplary waveform (blue) for IBM [PITH_FULL_IMAGE:figures/full_fig_p005_3.png]
Figure 4
Figure 4. Figure 4: FIG. 4. A4R waveform, which consists of four segments: ring [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. End-to-end workflow for machine learning-based qubit readout with QICK and hls4ml

    quant-ph 2025-01 conditional novelty 5.0 of 10

    A quantized neural network deployed on QICK's FPGA performs single-transmon readout at 96% fidelity, 32 ns inference latency, and under 16% LUT overhead, matching classical thresholding methods.

Reference graph

Works this paper leans on

49 extracted references · 30 canonical work pages · cited by 1 Pith paper

  1. [1]

    IBM Quantum, https://quantum.ibm.com/ (2021)

  2. [2]

    V. V. Sivak, A. Eickbusch, B. Royer, S. Singh, I. Tsiout- sios, S. Ganjam, A. Miano, B. L. Brock, A. Z. Ding, L. Frunzio, et al. , Real-time quantum error correction beyond break-even, Nature 616, 50–55 (2023)

  3. [3]

    Sundaresan, T

    N. Sundaresan, T. J. Yoder, Y. Kim, M. Li, E. H. Chen, G. Harper, T. Thorbeck, A. W. Cross, A. D. C´ orcoles, and M. Takita, Demonstrating multi-round subsystem quantum error correction using matching and maximum likelihood decoders, Nature Communications 14, 2852 (2023)

  4. [4]

    A. D. C´ orcoles, M. Takita, K. Inoue, S. Lekuch, Z. K. Minev, J. M. Chow, and J. M. Gambetta, Exploiting dynamic quantum circuits in a quantum algorithm with superconducting qubits, Phys. Rev. Lett. 127, 100501 (2021)

  5. [5]

    A. P. M. Place, L. V. H. Rodgers, P. Mundada, B. M. Smitham, M. Fitzpatrick, Z. Leng, A. Premkumar, J. Bryon, A. Vrajitoarea, S. Sussman, et al. , New mate- rial platform for superconducting transmon qubits with coherence times exceeding 0.3 milliseconds, Nature Com- munications 12, 10.1038/s41467-021-22030-5 (2021)

  6. [6]

    C. Wang, X. Li, H. Xu, Z. Li, J. Wang, Z. Yang, Z. Mi, X. Liang, T. Su, C. Yang, et al. , Towards practical quantum computers: transmon qubit with a lifetime ap- proaching 0.5 milliseconds, npj Quantum Information 8, 10.1038/s41534-021-00510-2 (2022)

  7. [7]

    Aumentado, Superconducting parametric amplifiers: The state of the art in josephson parametric amplifiers, IEEE Microwave Magazine 21, 45 (2020)

    J. Aumentado, Superconducting parametric amplifiers: The state of the art in josephson parametric amplifiers, IEEE Microwave Magazine 21, 45 (2020)

  8. [8]

    Vijay, D

    R. Vijay, D. H. Slichter, and I. Siddiqi, Observation of quantum jumps in a superconducting artificial atom, Phys. Rev. Lett. 106, 110502 (2011)

Show all 49 references
  1. [9]

    Ho Eom, P

    B. Ho Eom, P. K. Day, H. G. LeDuc, and J. Zmuidzinas, A wideband, low-noise superconducting amplifier with high dynamic range, Nature Physics 8, 623–627 (2012)

  2. [10]

    M. D. Reed, B. R. Johnson, A. A. Houck, L. DiCarlo, J. M. Chow, D. I. Schuster, L. Frunzio, and R. J. Schoelkopf, Fast reset and suppressing spontaneous emis- sion of a superconducting qubit, Applied Physics Letters 96 (2010)

  3. [11]

    Walter, P

    T. Walter, P. Kurpiers, S. Gasparinetti, P. Mag- nard, A. Potoˇ cnik, Y. Salath´ e, M. Pechal, M. Mondal, M. Oppliger, C. Eichler, and A. Wallraff, Rapid high- fidelity single-shot dispersive readout of superconducting qubits, Phys. Rev. Appl. 7, 054020 (2017)

  4. [12]

    Sunada, S

    Y. Sunada, S. Kono, J. Ilves, S. Tamate, T. Sugiyama, Y. Tabuchi, and Y. Nakamura, Fast readout and reset of a superconducting qubit coupled to a resonator with an intrinsic purcell filter, Phys. Rev. Appl. 17, 044016 (2022)

  5. [13]

    Dassonneville, T

    R. Dassonneville, T. Ramos, V. Milchakov, L. Planat, E. Dumur, F. Foroughi, J. Puertas, S. Leger, K. Bharad- waj, J. Delaforce, et al., Fast high-fidelity quantum non- demolition qubit readout via a nonperturbative cross- kerr coupling, Phys. Rev. X 10, 011045 (2020)

  6. [14]

    Swiadek, R

    F. Swiadek, R. Shillito, P. Magnard, A. Remm, C. Hellings, N. Lacroix, Q. Ficheux, D. C. Zanuz, G. J. Norris, A. Blais, et al. , Enhancing dispersive readout of superconducting qubits through dynamic control of the dispersive shift: Experiment and theory, PRX Quantum 5, 10.110...

  7. [15]

    Gambetta, W

    J. Gambetta, W. A. Braff, A. Wallraff, S. M. Girvin, and R. J. Schoelkopf, Protocols for optimal readout of qubits using a continuous quantum nondemolition mea- surement, Phys. Rev. A 76, 012325 (2007)

  8. [16]

    C. A. Ryan, B. R. Johnson, J. M. Gambetta, J. M. Chow, M. P. da Silva, O. E. Dial, and T. A. Ohki, Tomography via correlation of noisy measurement records, Phys. Rev. A 91, 022118 (2015)

  9. [17]

    C. C. Bultink, M. A. Rol, T. E. O’Brien, X. Fu, B. C. S. Dikken, C. Dickel, R. F. L. Vermeulen, J. C. de Sterke, A. Bruno, R. N. Schouten, and L. DiCarlo, Active res- onator reset in the nonlinear dispersive regime of circuit qed, Phys. Rev. Appl. 6, 034008 (2016)

  10. [18]

    D. T. McClure, H. Paik, L. S. Bishop, M. Steffen, J. M. Chow, and J. M. Gambetta, Rapid driven reset of a qubit readout resonator, Phys. Rev. Appl. 5, 011001 (2016)

  11. [19]

    Bengtsson, A

    A. Bengtsson, A. Opremcak, M. Khezri, D. Sank, A. Bourassa, K. J. Satzinger, S. Hong, C. Erickson, B. J. Lester, K. C. Miao, et al. , Model-based optimization of superconducting qubit readout, Phys. Rev. Lett. 132, 11 100603 (2024)

  12. [20]

    R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed. (The MIT Press, 2018)

  13. [21]

    S. M. Kakade, A natural policy gradient, in Advances in Neural Information Processing Systems , Vol. 14, edited by T. Dietterich, S. Becker, and Z. Ghahramani (MIT Press, 2001)

  14. [22]

    Fran¸ cois-Lavet, P

    V. Fran¸ cois-Lavet, P. Henderson, R. Islam, M. G. Belle- mare, and J. Pineau, An introduction to deep reinforce- ment learning, Foundations and Trends ® in Machine Learning 11, 219–354 (2018)

  15. [23]

    Schulman, S

    J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel, Trust region policy optimization (2015)

  16. [24]

    AnikenC, Github - anikenc/rl-meets-qubit-readout: Repository for demonstration of enhanced qubit readout via reinforcement learning (2024)

  17. [25]

    Blais, A

    A. Blais, A. L. Grimsmo, S. M. Girvin, and A. Wallraff, Circuit quantum electrodynamics, Rev. Mod. Phys. 93, 025005 (2021)

  18. [26]

    M. F. Dumas, B. Groleau-Par´ e, A. McDonald, M. H. Mu˜ noz Arias, C. Lled´ o, B. D’Anjou, and A. Blais, Measurement-induced transmon ionization, Phys. Rev. X 14, 041023 (2024)

  19. [27]

    Hoffer, Superconducting qubit readout pulse optimiza- tion using deep reinforcement learning (2021)

    C. Hoffer, Superconducting qubit readout pulse optimiza- tion using deep reinforcement learning (2021)

  20. [28]

    Lienhard, Machine learning assisted superconducting qubit readout (2021)

    B. Lienhard, Machine learning assisted superconducting qubit readout (2021)

  21. [29]

    J. A. Nelder and R. Mead, A simplex method for function minimization, Computer Journal 7, 308 (1965)

  22. [30]

    Kennedy and R

    J. Kennedy and R. Eberhart, Particle swarm optimiza- tion, in Proceedings of ICNN’95 - International Confer- ence on Neural Networks , ICNN-95, Vol. 4 (IEEE) p. 1942–1948

  23. [31]

    Cheng, X.-J

    X. Cheng, X.-J. Lu, Y.-N. Liu, and S. Kuang, Compar- ison of differential evolution, particle swarm optimiza- tion, quantum-behaved particle swarm optimization, and quantum evolutionary algorithm for preparation of quan- tum states, Chinese Physics B 32, 020202 (2023)

  24. [32]

    Y. Baum, M. Amico, S. Howell, M. Hush, M. Liuzzi, P. Mundada, T. Merkh, A. R. R. Carvalho, and M. J. Biercuk, Experimental deep reinforcement learning for error-robust gate-set design on a superconducting quan- tum computer, PRX Quantum 2, 040324 (2021)

  25. [33]

    J. Olle, R. Zen, M. Puviani, and F. Marquardt, Si- multaneous discovery of quantum error correction codes and encoders with a noise-aware reinforcement learning agent, npj Quantum Information 10, 10.1038/s41534- 024-00920-y (2024)

  26. [34]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, Proximal policy optimization algorithms (2017), arXiv:1707.06347

  27. [35]

    Konda and J

    V. Konda and J. Tsitsiklis, Actor-critic algorithms, in Advances in Neural Information Processing Systems , Vol. 12, edited by S. Solla, T. Leen, and K. M¨ uller (MIT Press, Cambridge, MA, USA, 1999)

  28. [36]

    Sowerby, Z

    H. Sowerby, Z. Zhou, and M. L. Littman, Designing re- wards for fast learning (2022)

  29. [37]

    Jeffrey, D

    E. Jeffrey, D. Sank, J. Y. Mutus, T. C. White, J. Kelly, R. Barends, Y. Chen, Z. Chen, B. Chiaro, A. Dunsworth, et al. , Fast accurate state measurement with supercon- ducting qubits, Phys. Rev. Lett. 112, 190504 (2014)

  30. [38]

    Proctor, M

    T. Proctor, M. Revelle, E. Nielsen, K. Rudinger, D. Lob- ser, P. Maunz, R. Blume-Kohout, and K. Young, Detect- ing and tracking drift in quantum information processors, Nature Communications 11, 10.1038/s41467-020-19074- 4 (2020)

  31. [39]

    Hazra, W

    S. Hazra, W. Dai, T. Connolly, P. D. Kurilovich, Z. Wang, L. Frunzio, and M. H. Devoret, Benchmark- ing the readout of a superconducting qubit for repeated measurements, Phys. Rev. Lett. 134, 100601 (2025)

  32. [40]

    P. A. Spring, L. Milanovic, Y. Sunada, S. Wang, A. F. van Loo, S. Tamate, and Y. Nakamura, Fast multiplexed superconducting qubit readout with intrinsic purcell fil- tering (2024), arXiv:2409.04967 [quant-ph]

  33. [41]

    Bradbury, R

    J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. Van- derPlas, S. Wanderman-Milne, and Q. Zhang, JAX: com- posable transformations of Python+NumPy programs (2018)

  34. [42]

    C. Lu, J. Kuba, A. Letcher, L. Metz, C. Schroeder de Witt, and J. Foerster, Discovered policy optimisation, Advances in Neural Information Processing Systems 35, 16455 (2022)

  35. [43]

    Huang, R

    S. Huang, R. F. J. Dossa, A. Raffin, A. Kanervisto, and W. Wang, The 37 implementation details of prox- imal policy optimization, in ICLR Blog Track (2022) https://iclr-blog-track.github.io/2022/03/25/ppo- implementation-details/

  36. [44]

    A. Bou, M. Bettini, S. Dittert, V. Kumar, S. Sodhani, X. Yang, G. De Fabritiis, and V. Moens, Torchrl: A data- driven decision-making library for pytorch (2023)

  37. [45]

    J. B. Curtis, I. Boettcher, J. T. Young, M. F. Maghrebi, H. Carmichael, A. V. Gorshkov, and M. Foss-Feig, Crit- ical theory for the breakdown of photon blockade, Phys. Rev. Res. 3, 023062 (2021)

  38. [46]

    D. I. Schuster, A. Wallraff, A. Blais, L. Frunzio, R. S. Huang, J. Majer, S. M. Girvin, and R. J. Schoelkopf, ac stark shift and dephasing of a superconducting qubit strongly coupled to a cavity field, Physical Review Let- ters 94, 10.1103/PhysRevLett.94.123602 (2005)

  39. [47]

    Gambetta, A

    J. Gambetta, A. Blais, D. I. Schuster, A. Wallraff, L. Frunzio, J. Majer, M. H. Devoret, S. M. Girvin, and R. J. Schoelkopf, Qubit-photon interactions in a cavity: Measurement-induced dephasing and number splitting, Physical Review A 74, 10.1103/PhysRevA.74.042318 (2006)

  40. [48]

    Kidger, On Neural Differential Equations , Ph.D

    P. Kidger, On Neural Differential Equations , Ph.D. the- sis, University of Oxford (2021)

  41. [49]

    Virtanen, R

    P. Virtanen, R. Gommers, T. E. Oliphant, M. Haber- land, T. Reddy, D. Cournapeau, E. Burovski, P. Peter- son, W. Weckesser, J. Bright, et al. , SciPy 1.0: Funda- mental Algorithms for Scientific Computing in Python, Nature Methods 17, 261 (2020)

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.