REVIEW 3 major objections 6 minor 39 references
Quantum Error Management in Practice: A Cross-Stack Benchmark
T0 review · 3 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Within one execution campaign on a 156-qubit IBM processor, managed error stacks beat IBM's built-in options on every sampling test and cut estimation error by 3.1–4.7x.
desk verdict Useful first same-hardware, same-workload benchmark of three commercial error-management stacks, but the single-run design and missing calibration data mean the cross-provider ranking is not yet stable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the benchmark protocol itself: a seeded, manifest-driven submit-and-collect workflow that runs the same logical circuits on the same processor through each provider's native interface, with IBM and Q-CTRL at 32,768 shots and QESEM at its precision-controlled 600-second budget. The workloads are chosen so correctness is unambiguous: Bernstein–Vazirani, quantum phase estimation, GHZ preparation, and randomized mirror circuits have exact ideal bitstrings, scored by exact-output probability $P_{\text{exact}}$ or valid-set probability $P_{\text{valid}}$, and the eight-layer transverse-field Ising circuit is scored against a machine-precision matrix-product-state reference for chain-averaged magnetization and correlator observables. The metrics—exact-output probability, valid-set probability, absolute error $\Delta$, and aggregate mean absolute error—carry the argument by making cross-provider accuracy comparable on a single axis.
What would settle it
Rerun the same benchmark suite many times per configuration across several calibration windows on the same processor, producing distributions of success probabilities and mean absolute error instead of single numbers; if IBM raw or TREX-plus-twirling ever matches Q-CTRL or QESEM within the spread, the claim that managed stacks dominate in this setting would be falsified. A cheaper check is to run ten QESEM jobs at the same precision target and see whether the 7.5–11.1x QPU-time cost and the 0.0188 aggregate MAE are stable.
Extended reading notes
Core claim
The paper's central claim is that, within this execution campaign, managed error suppression and mitigation outperformed IBM's built-in error-handling options on both sampling and estimation tasks, with a clear accuracy–time tradeoff. On the Sampler track, Q-CTRL's Performance Management function returned the highest exact-output or valid-set probability in all twelve instances, including 76.5% exact 100-qubit mirror-circuit success versus 9.1% raw, and 12.7% exact phase-estimation success at 30 counting qubits where raw and twirled IBM execution returned zero. On the Estimator track, aggregate mean absolute error against the matrix-product-state reference was 0.0883 for IBM raw, 0.0807 for IBM TREX plus twirling, 0.0285 for Q-CTRL, and 0.0188 for QESEM, so the managed stacks reduced aggregate error by factors of 3.10 and 4.70. QESEM achieved the lowest error while using 211–311 seconds of reported QPU time per job versus 28 seconds for Q-CTRL. The paper also finds that IBM's TREX-plus-twirling configuration improved the magnetization but systematically worsened the correlator, showing that a mitigation that helps one observable can hurt another.
Load-bearing premise
The entire ranking rests on a single execution job per provider configuration at each size, so if calibration drift or job-to-job variability between provider submissions is larger than the measured accuracy differences, the cross-provider ordering could change.
Editorial extensions
If this is right
- Practitioners running sampling workloads who need the correct bitstring, not just an averaged expectation value, can expect managed suppression pipelines like Q-CTRL to raise success probabilities by large factors, especially on deep or wide circuits.
- For expectation-value workloads, managed stacks can cut error by roughly 3 to 5 times relative to raw IBM execution, and the cheaper managed option already operates within the same QPU-time order as IBM's built-in configurations.
- QESEM's higher accuracy comes at 7.5 to 11.1 times the reported QPU time of Q-CTRL, so accuracy gains have a real resource cost that must be weighed against job budgets and queue time.
- Measurement twirling alone is unlikely to improve exact-bitstring metrics, because it symmetrizes readout errors rather than reducing the total error rate; its value is confined to expectation-value tasks.
- The seeded workload suite and manifest-driven protocol are a reusable template for like-for-like comparisons of commercial error-management layers, including on early fault-tolerant hardware where mitigation will still sit on top of logical qubits.
Reading between the lines
- If the accuracy ordering replicates across calibration windows, a cost-normalized ranking would likely change: QESEM's 4.70x error reduction over raw execution is only 1.5 times better than Q-CTRL's 3.10x, yet costs 7.5–11.1x more QPU time, so the paper's raw accuracy ranking is not a recommendation for every budget.
- The Sampler results suggest that for bitstring-answer tasks, suppression is the dominant lever and expectation-value mitigation cannot rescue individual shots; a testable extension would apply distribution-level readout mitigation on top of raw counts and measure whether exact-output success improves.
- The observable-dependent behavior of TREX-plus-twirling argues that single-observable benchmarks can mislead; scoring multiple observables from the same circuit, as done here with magnetization and correlator, should become the default for mitigation comparisons.
- The symmetry argument that excludes longitudinal magnetization from the Ising benchmark is a model for designing fair workloads: observables pinned to zero by symmetry reward noise-damped answers and test nothing about accuracy.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports a single execution campaign on ibm_pittsburgh, a 156-qubit Heron r3 processor, comparing IBM Qiskit Runtime, Q-CTRL Performance Management, and Qedma QESEM. On the Sampler track, Bernstein–Vazirani, quantum phase estimation, GHZ preparation, and randomized mirror circuits are scored by exact-output or valid-set probability. On the Estimator track, transverse-field Ising observables at 25, 50, and 75 qubits are scored against an exact matrix-product-state reference, with IBM raw, IBM TREX plus twirling, Q-CTRL, and QESEM configurations. The paper reports that Q-CTRL achieved the highest Sampler success on every tested instance, and that Q-CTRL and QESEM reduced aggregate mean absolute error relative to IBM raw by factors of 3.10 and 4.70 respectively, with QESEM using 7.5–11.1 times the reported QPU time of Q-CTRL. The authors frame these as results 'within this execution campaign' and list limitations in Section 4.
Significance. If the reported accuracy gaps are reproducible across calibration cycles, this is a valuable independent comparison of commercial error-management stacks on identical workloads and hardware. The paper's strengths include exact ideal outputs for the Sampler families, an MPS reference computed at machine precision, a seeded manifest-driven workflow with recorded job identifiers, stated shot budgets, and explicit qualifications about provider-specific uncertainty fields. However, the single-job design means the cross-provider ordering is currently an observation from one time window, not a statistically supported benchmark; the paper also contains an unresolved inconsistency in the description of the IBM resilience-level setting. These issues affect the central comparative claims and the reproducibility of the IBM baseline.
major comments (3)
- [Section 4; Sections 2.1 and 2.2.1] The central claim—Q-CTRL winning every Sampler instance and the aggregate MAE reductions of 3.10 and 4.70—rests on a single execution job per configuration, as stated in Section 4: 'Each provider instance condition was represented by one execution job.' The shot-level binomial error bounds quoted in Section 2.2.1 address finite sampling only; they do not bound job-to-job hardware variability, calibration drift between provider submissions, or compilation-dependent variation. Without per-window calibration metrics (e.g., median two-qubit error, readout error, T1/T2) or interleaved repeated runs, the observed ordering could change if the Q-CTRL or QESEM jobs landed in a better device state. This is load-bearing for the paper's practical guidance. To make the claim robust, the authors should either add calibration data for each execution window and show the gaps exceed plausible drift, or rerun a subset of instances interleaved across calibration windows and report variance; alternatively, the abstract and conclusions should be downgraded to a single-campaign observation without general performance ordering.
- [Section 2.3.3; Table 3; Section 4] The IBM TREX+twirling arm is described inconsistently. Section 2.3.3 states 'The second configuration used resilience level 1 together with explicitly enabled gate twirling,' while Table 3 labels the same configuration as 'resilience level 2; TREX measurement mitigation on; gate + measurement twirling on; ZNE/PEC off.' Section 4 then refers to 'the tested IBM resilience-level-2 configuration.' Additionally, Section 2.3.3 says ZNE was configured at resilience level 2 but not executed, which conflicts with Table 3's 'ZNE/PEC off.' This ambiguity undermines the reproducibility of the IBM baseline arm and could mislead readers about exactly which Qiskit Runtime options were tested. Please reconcile the resilience-level label and specify precisely which options were enabled, which were disabled, and which were configured but not executed.
- [Section 2.1; Section 3.3; Table 6] The Estimator comparison mixes operating points: IBM and Q-CTRL used matched 32,768-shot budgets, while QESEM used a provider-native precision target of 0.1 and a 600-second QPU cap. The paper acknowledges this, but the aggregate error-reduction factors in the abstract and Section 3.3 are not resource-normalized. A reader could interpret 'QESEM reduced aggregate error by a factor of 4.70' as an accuracy gain achieved at comparable cost. Please state explicitly in the abstract and results that the QESEM reduction is obtained at 7.5–11.1 times the QPU time of Q-CTRL and that no normalization for QPU time or monetary cost is attempted.
minor comments (6)
- [Data availability] The paper says the benchmark notebooks, manifests, seeds, and scripts 'are available from the corresponding author on reasonable request.' Given the explicit reproducibility claim in the abstract, please deposit these artifacts in a permanent public repository with a DOI.
- [Section 4] The sentence 'The comparison therefore includes the tested IBM resilience-level-2 configuration with explicitly enabled gate twirling, but does not include IBM ZNE or PEC' is confusing given the Table 3 label; please rephrase to match the corrected resilience-level terminology.
- [Table 4 note] The one-sided 95% upper confidence bound of approximately 0.0091% for zero successes in 32,768 shots is plausible, but the text should state the method used (e.g., Clopper–Pearson or Wilson).
- [Figures 11 and 13] The captions of Figures 11 and 13 are terse; please specify that the plotted quantity is the normalized success score of Eq. (6) for the corresponding observable.
- [Section 1] The paragraph beginning 'The crucial limitation is that mitigation reconstructs expectation values of observables' contains a repeated phrase ('The crucial limitation is') and would benefit from editing to a single sentence.
- [References] Several references are to vendor documentation with access dates (Refs. 22, 23, 25, 28); please include consistent access dates and version identifiers where available.
Circularity Check
No significant circularity: the benchmark scores provider outputs against external ideal-bitstring and MPS references, with no fitted parameters and no derivation that assumes its own conclusions.
full rationale
The paper's central claims are empirical comparisons of provider outputs against external references, not derivations from assumptions containing those claims. Sampler success is scored against ideal single bitstrings or the two-string GHZ support (Eqs. 1 and 2), and Estimator errors are measured against an independent matrix-product-state simulation with truncation threshold 1e-16 (Section 2.3.1). No parameter is fitted to any provider's output, and the aggregate MAE reduction factors in Table 7 are arithmetic summaries of measured per-case absolute errors rather than predictions implied by the scoring definitions. The only author self-citation, Ref. 27 in Section 2.2.5, notes a prior use of the mirror-circuit inversion strategy, but that strategy is itself attributed to the external mirror-circuit benchmarking literature (Ref. 35) and is not load-bearing for any result. The paper's stated limitations, including single-job representation, provider-reported QPU times, and provider-specific stds fields, affect statistical generality and measurement provenance but do not make any step circular. No step in the paper reduces to its own input by construction.
Assumptions & free parameters
assumptions (4)
- domain assumption The MPS reference with singular-value truncation 1e-16 is numerically exact for the 1D eight-layer TFIM circuits at sizes 25, 50, and 75.
- domain assumption Provider-reported QPU times are accurate and comparable across providers.
- domain assumption One execution job per configuration per size is representative of provider performance.
- domain assumption Treating each provider via its native interface (transpiled ISA for IBM, abstract circuits for Q-CTRL and QESEM) is a fair like-for-like comparison.
Cite this review
Pith. "Pith review of Quantum Error Management in Practice: A Cross-Stack Benchmark." pith.science (2026). https://pith.science/paper/BWBK5LIT
@misc{pith2026260805202,
author = {Pith},
title = {Pith review of: Quantum Error Management in Practice: A Cross-Stack Benchmark},
year = {2026},
howpublished = {\url{https://pith.science/paper/BWBK5LIT}},
note = {Machine review of arXiv:2608.05202}
}
read the original abstract
Quantum processors have crossed the one-hundred-qubit mark, but noise continues to limit circuit performance, while full quantum error correction remains too costly for routine use. Error suppression and mitigation therefore play an important role in extracting value from current hardware, yet independent comparisons of commercial solutions on identical workloads and devices remain scarce. We benchmark IBM Qiskit Runtime, Q-CTRL Performance Management, and Qedma QESEM on IBM Pittsburgh, a 156-qubit IBM Quantum Heron r3 processor. For Sampler workloads, we run Bernstein-Vazirani, quantum phase estimation, GHZ-state preparation, and randomized mirror circuits with up to 100 measured qubits, comparing raw execution, IBM measurement twirling, and Q-CTRL. For Estimator workloads, we measure chain-averaged magnetization and correlation observables of an eight-layer transverse-field Ising circuit at 25, 50, and 75 qubits against an exact matrix-product-state reference, comparing IBM raw execution, IBM TREX plus twirling, Q-CTRL, and QESEM. Q-CTRL produced the best results on the three structured Sampler workloads while keeping reported QPU times within the same order as the IBM configurations. Across six Ising observable and system-size cases, aggregate mean absolute error was 0.0883 for IBM raw execution, 0.0807 for IBM TREX plus twirling, 0.0285 for Q-CTRL, and 0.0188 for QESEM. Relative to raw execution, Q-CTRL and QESEM reduced aggregate error by factors of 3.10 and 4.70, respectively, while QESEM used 7.5 to 11.1 times the reported QPU time of Q-CTRL. These results show that managed error suppression and mitigation can substantially improve current hardware performance, but with distinct accuracy and execution-time tradeoffs.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Quantum computing in the NISQ era and beyond,
J. Preskill, “Quantum computing in the NISQ era and beyond,”Quantum2, 79 (2018)
work page 2018
-
[2]
Quantum computing optimization, an intro- duction,
A. Giraldo-Quintero, D. Sierra-Sosa, I. Imam, and A. Elmaghraby, “Quantum computing optimization, an intro- duction,”IEEE Access, early access (2026)
work page 2026
-
[3]
TensorFlow Quantum: Impacts of quantum state preparation on quantum machine learning performance,
D. Sierra-Sosa, M. Telahun, and A. Elmaghraby, “TensorFlow Quantum: Impacts of quantum state preparation on quantum machine learning performance,”IEEE Access8, 215246–215255 (2020). 19 Quantum Error Management in PracticePREPRINT
work page 2020
-
[4]
Variational quantum classifier for binary classification: Real vs synthetic dataset,
D. Maheshwari, D. Sierra-Sosa, and B. Garcia-Zapirain, “Variational quantum classifier for binary classification: Real vs synthetic dataset,”IEEE Access10, 3705–3715 (2022)
work page 2022
-
[5]
Scheme for reducing decoherence in quantum computer memory,
P. W. Shor, “Scheme for reducing decoherence in quantum computer memory,”Phys. Rev. A52, R2493 (1995)
work page 1995
-
[6]
Error correcting codes in quantum theory,
A. M. Steane, “Error correcting codes in quantum theory,”Phys. Rev. Lett.77, 793 (1996)
work page 1996
-
[7]
Surface codes: Towards practical large-scale quantum computation,
A. G. Fowler, M. Mariantoni, J. M. Martinis, and A. N. Cleland, “Surface codes: Towards practical large-scale quantum computation,”Phys. Rev. A86, 032324 (2012)
work page 2012
-
[8]
Suppressing quantum errors by scaling a surface code logical qubit,
Google Quantum AI, “Suppressing quantum errors by scaling a surface code logical qubit,”Nature614, 676–681 (2023)
work page 2023
Show all 39 references
-
[9]
Quantum error correction below the surface code threshold,
Google Quantum AI and Collaborators, “Quantum error correction below the surface code threshold,”Nature 638, 920–926 (2025)
2025
-
[10]
High-threshold and low-overhead fault-tolerant quantum memory,
S. Bravyi, A. W. Cross, J. M. Gambetta, D. Maslov, P. Rall, and T. J. Yoder, “High-threshold and low-overhead fault-tolerant quantum memory,”Nature627, 778–782 (2024)
2024
-
[11]
Dynamical decoupling of open quantum systems,
L. Viola, E. Knill, and S. Lloyd, “Dynamical decoupling of open quantum systems,”Phys. Rev. Lett.82, 2417 (1999)
1999
-
[12]
Noise tailoring for scalable quantum computation via randomized compiling,
J. J. Wallman and J. Emerson, “Noise tailoring for scalable quantum computation via randomized compiling,” Phys. Rev. A94, 052325 (2016)
2016
-
[13]
Error mitigation for short-depth quantum circuits,
K. Temme, S. Bravyi, and J. M. Gambetta, “Error mitigation for short-depth quantum circuits,”Phys. Rev. Lett. 119, 180509 (2017)
2017
-
[14]
Efficient variational quantum simulator incorporating active error minimization,
Y . Li and S. C. Benjamin, “Efficient variational quantum simulator incorporating active error minimization,” Phys. Rev. X7, 021050 (2017)
2017
-
[15]
Error mitigation extends the computational reach of a noisy quantum processor,
A. Kandala, K. Temme, A. D. Córcoles, A. Mezzacapo, J. M. Chow, and J. M. Gambetta, “Error mitigation extends the computational reach of a noisy quantum processor,”Nature567, 491–495 (2019)
2019
-
[16]
Probabilistic error cancellation with sparse Pauli– Lindblad models on noisy quantum processors,
E. van den Berg, Z. K. Minev, A. Kandala, and K. Temme, “Probabilistic error cancellation with sparse Pauli– Lindblad models on noisy quantum processors,”Nat. Phys.19, 1116–1121 (2023)
2023
-
[17]
Model-free readout-error mitigation for quantum expectation values,
E. van den Berg, Z. K. Minev, and K. Temme, “Model-free readout-error mitigation for quantum expectation values,”Phys. Rev. A105, 032620 (2022)
2022
-
[18]
Quantum error mitigation,
Z. Cai, R. Babbush, S. C. Benjamin, S. Endo, W. J. Huggins, Y . Li, J. R. McClean, and T. E. O’Brien, “Quantum error mitigation,”Rev. Mod. Phys.95, 045005 (2023)
2023
-
[19]
Evidence for the utility of quantum computing before fault tolerance,
Y . Kim, A. Eddins, S. Anand, K. X. Wei, E. van den Berg, S. Rosenblatt, H. Nayfeh, Y . Wu, M. Zaletel, K. Temme, and A. Kandala, “Evidence for the utility of quantum computing before fault tolerance,”Nature 618, 500–505 (2023)
2023
-
[20]
On the importance of error mitigation for quantum computation,
D. Aharonov, O. Alberton, I. Arad, et al. (Qedma Quantum Computing), “On the importance of error mitigation for quantum computation,” arXiv:2503.17243 (2025)
2025 arXiv
-
[21]
Quantum computing with Qiskit,
A. Javadi-Abhari, M. Treinish, K. Krsulich, C. J. Wood, J. Lishman, J. Gacon, S. Martiel, P. D. Nation, L. S. Bishop, A. W. Cross, B. R. Johnson, and J. M. Gambetta, “Quantum computing with Qiskit,” arXiv:2405.08810 (2024)
2024 arXiv
-
[22]
Error mitigation and suppression techniques,
IBM Quantum, “Error mitigation and suppression techniques,” IBM Quantum Documentation,https:// quantum.cloud.ibm.com/docs/guides/error-mitigation-and-suppression-techniques(accessed 2026)
2026
-
[23]
Introduction to Qiskit Functions,
IBM Quantum, “Introduction to Qiskit Functions,” IBM Quantum Documentation,https://quantum.cloud. ibm.com/docs/guides/functions(accessed 2026)
2026
-
[24]
Experimental benchmarking of an automated deterministic error-suppression work- flow for quantum algorithms,
P. S. Mundada, A. Barbosa, S. Maity, Y . Wang, T. M. Stace, T. Merkh, F. Nielson, A. R. R. Carvalho, M. Hush, M. J. Biercuk, and Y . Baum, “Experimental benchmarking of an automated deterministic error-suppression work- flow for quantum algorithms,”Phys. Rev. Applied20, 024034 (2023)
2023
-
[25]
Performance Management: A Qiskit Function by Q-CTRL Fire Opal,
Q-CTRL, “Performance Management: A Qiskit Function by Q-CTRL Fire Opal,” IBM Quantum Documen- tation,https://quantum.cloud.ibm.com/docs/guides/q-ctrl-performance-management(accessed 2026)
2026
-
[26]
Reliable high-accuracy error mitigation for utility-scale quantum circuits,
D. Aharonov et al., “Reliable high-accuracy error mitigation for utility-scale quantum circuits,” arXiv:2508.10997 (2025)
2025 arXiv
-
[27]
A teleportation protocol variant for single-QPU benchmarking,
C. Márquez, D. Sierra-Sosa, and K. Garcés, “A teleportation protocol variant for single-QPU benchmarking,” IEEE Access13, 209266–209281 (2025). 20 Quantum Error Management in PracticePREPRINT
2025
-
[28]
QESEM: A Qiskit Function by Qedma,
Qedma, “QESEM: A Qiskit Function by Qedma,” IBM Quantum Documentation,https://quantum.cloud. ibm.com/docs/guides/qedma-qesem(accessed 2026)
2026
-
[29]
Welcome ibm_pittsburgh, goodbye ibm_sherbrooke,
IBM Quantum, “Welcome ibm_pittsburgh, goodbye ibm_sherbrooke,” IBM Quantum Platform prod- uct update (July 31, 2025),https://quantum.cloud.ibm.com/announcements/en/product-updates/ 2025-07-31-pittsburgh-sherbrooke
2025
-
[30]
Quantum complexity theory,
E. Bernstein and U. Vazirani, “Quantum complexity theory,”SIAM J. Comput.26, 1411–1473 (1997)
1997
-
[31]
Quantum measurements and the Abelian stabilizer problem,
A. Yu. Kitaev, “Quantum measurements and the Abelian stabilizer problem,” arXiv:quant-ph/9511026 (1995)
1995 arXiv
-
[32]
M. A. Nielsen and I. L. Chuang,Quantum Computation and Quantum Information, 10th anniversary ed. (Cam- bridge University Press, Cambridge, 2010)
2010
-
[33]
Going beyond Bell’s theorem,
D. M. Greenberger, M. A. Horne, and A. Zeilinger, “Going beyond Bell’s theorem,” inBell’s Theorem, Quantum Theory and Conceptions of the Universe, edited by M. Kafatos (Kluwer, Dordrecht, 1989), pp. 69–72
1989
-
[34]
Generation and verification of 27-qubit Greenberger–Horne–Zeilinger states in a superconducting quantum computer,
G. J. Mooney, G. A. L. White, C. D. Hill, and L. C. L. Hollenberg, “Generation and verification of 27-qubit Greenberger–Horne–Zeilinger states in a superconducting quantum computer,”J. Phys. Commun.5, 095004 (2021)
2021
-
[35]
Measuring the capabilities of quantum computers,
T. Proctor, K. Rudinger, K. Young, E. Nielsen, and R. Blume-Kohout, “Measuring the capabilities of quantum computers,”Nat. Phys.18, 75–79 (2022)
2022
-
[36]
The one-dimensional Ising model with a transverse field,
P. Pfeuty, “The one-dimensional Ising model with a transverse field,”Ann. Phys. (N.Y.)57, 79–90 (1970)
1970
-
[37]
Sachdev,Quantum Phase Transitions, 2nd ed
S. Sachdev,Quantum Phase Transitions, 2nd ed. (Cambridge University Press, Cambridge, 2011)
2011
-
[38]
Efficient classical simulation of slightly entangled quantum computations,
G. Vidal, “Efficient classical simulation of slightly entangled quantum computations,”Phys. Rev. Lett.91, 147902 (2003)
2003
-
[39]
The density-matrix renormalization group in the age of matrix product states,
U. Schollwöck, “The density-matrix renormalization group in the age of matrix product states,”Ann. Phys. (N.Y.) 326, 96–192 (2011). 21
2011
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.