REVIEW 3 major objections 5 minor 22 references
Dynamical codes for hardware with noisy readouts
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Repeating measurements in a dynamical quantum error-correcting code shrinks its spacetime cost when readout noise dominates, but not otherwise.
desk verdict Useful, honest numerical study of DCCC schedule tailoring; the teraquop-volume ordering rests on extrapolation and should be read as provisional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the measurement schedule of a dynamically condensed colour code, specified by how many consecutive rounds of XX, YY, and ZZ edge measurements are performed, giving XaYbZc or XaZb codes. Repeating measurements is schedule-induced gauge-fixing: it creates two-measurement edge detectors that catch measurement errors directly, at the cost of lengthening the face detectors and worsening timelike distances for data-qubit errors. The paper's performance metric is the teraquop volume, the product of qubit count and measurement rounds needed to reach a $10^{-12}$ logical error rate. The comparison tool is the decoding graph, where measurement errors that violate four detectors become hyperedges; belief matching pre-processes these hyperedge probabilities before minimum-weight matching, recovering information that ordinary MWPM throws away.
What would settle it
Run the same memory and stability experiments at substantially larger distances, say distance 24 or 32 for phenomenological noise, and check whether logical error rates continue the exponential decay that the line fits assume. A second check is to simulate the X1Z1 code with planar boundaries and compare its logical error rate to the toric result; if planar performance differs markedly, the volume rankings built on toric simulations would need revisiting.
Extended reading notes
Core claim
The paper's central claim is that the optimal DCCC measurement schedule is not intrinsic to the code but is co-determined by noise bias and decoder. In all three phenomenological noise models studied, the X1Y1Z1 code under belief matching has the smallest teraquop volume, while the same code under MWPM has the largest; the effect grows with measurement bias. Repeating measurements, a form of schedule-induced gauge-fixing that creates two-measurement edge detectors, improves the teraquop volumes of both the XaZb and XaYbZc code families under both decoders when the noise is measurement-biased, but is negligible otherwise, contrary to the authors' initial expectations. At the circuit level, the superconducting-inspired noise bias is not strong enough to make repetition worthwhile, while under entangling-measurement noise the X2Z2 code wins with MWPM. Across most of the parameter sweep, performance differences come primarily from the number of measurement rounds required rather than the number of qubits, which is why the volume metric matters.
Load-bearing premise
The teraquop volumes are extrapolated from logical error rates at small code distances down to $10^{-12}$ by assuming exponential decay in one code dimension and linear dependence on the others; if error rates bend at larger sizes, the reported volumes and rankings could change. The simulations also assume a torus, and planar performance for the X1Z1 code has not been verified.
Editorial extensions
If this is right
- Lattice-surgery resource estimates that use only the teraquop footprint will miss most of the difference between codes, since the gaps come primarily from the number of measurement rounds.
- Hardware with measurement-dominated noise should use schedules with repeated measurements; hardware with unbiased or Z-biased noise gains little from repetition.
- Decoder choice should be made jointly with code choice, because belief matching can turn a worst-performing code into the best-performing one under the same noise model.
- Under entangling-measurement noise and MWPM, the X2Z2 schedule wins, showing that the optimal schedule also depends on how the multi-qubit measurement is physically implemented.
Reading between the lines
- For a readout-limited device, the practical recipe suggested by these results is a measurement-repeated X1Y1Z1 schedule decoded with belief matching; the paper does not spell this out as a single recommendation, but it follows from the per-bias best-code tables.
- The teraquop volume metric could be used to benchmark other spacetime codes against DCCCs, but this paper only compares codes within the DCCC family.
- At Z bias beyond 16, asymmetric repetition schedules that spend more time in detectors sensitive to Z errors may begin to pay off; the paper's parameter sweep stops too early to see this.
- The reported ranking may be sensitive to the assumption that timelike E and M errors occur equally often in a lattice-surgery computation; if one type dominates, the optimal height differs from the averaged value used here.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies dynamically condensed colour codes (DCCCs) under noise models with tunable measurement bias and Z bias, and introduces the teraquop volume as a spacetime generalization of the teraquop footprint. Using Stim/Sinter with PyMatching (MWPM) and BeliefMatching, it simulates XaYbZc and XaZb codes on a torus for phenomenological, standard depolarizing, superconducting-inspired, and entangling-measurement noise. It reports that (i) MWPM favors XaZb schedules while belief matching favors XaYbZc schedules, (ii) differences in volume are dominated by the number of measurement rounds rather than qubit count, and (iii) repeating measurements helps mainly under strong measurement bias. The paper also provides analytic distance and hyperedge-likelihood tables and makes the simulation code publicly available.
Significance. If the reported teraquop-volume rankings are correct, the paper gives a concrete, decoder- and hardware-aware rule for choosing DCCC schedules, and the teraquop volume is a useful metric for comparing spacetime overhead. The strengths are the use of a standard, externally benchmarked simulation pipeline (Stim/Sinter/PyMatching/BeliefMatching) with explicit noise models and error bars; the public code for full reproducibility; and the transparent analytic tables (Tables 3, 4, and 6) that separate distance effects from decoder-dependent hyperedge effects. The central caveat is that the headline volumes and orderings are extrapolated from small simulated distances, so the quantitative ranking is less secure than the qualitative trends. The finding that volume differences come predominantly from measurement rounds rather than footprint is an important caution for metric choice in QEC resource estimation.
major comments (3)
- [Section 3.1] The headline volumes and rankings are extrapolated, not directly simulated. The fits use distances d=4,8,12,16 (d=2,4,6,8 for EM noise) and assume log p_L is linear in d with linear dependence on the other two dimensions, then extend to 10^-12. Figure 16 validates the linear-in-nM and linear-in-h dependence only for E memory experiments under MWPM for the X1Z1 and X1Y1Z1 codes; no equivalent check is shown for belief matching or for repeated codes. The belief-matching fits use 10^4 shots per point (Appendix A), so at the largest d a run may have zero or very few logical errors, making the fitted slope and intercept carry large relative uncertainty. Since the ranking differences in Key points 4.5 and 5.7 are factors of a few in the teraquop height, a modest change in slope or a bend at d>16 could reorder the codes. Please add data at larger distances or provide a bounded extrapolation analysis, report the number of logical errors per point for the belief-matching fits, and state whether the linear-dependence assumption has been checked for the repeated schedules and for belief matching.
- [Definition 3.1] The definition requires the probability of any spacelike or timelike logical error to be at most 10^-12, but the estimation procedure minimizes each of nE, nM, hE, and hM separately from four different experiments and then forms the product nE * nM * (hE+hM)/2. The text does not give a union bound or an independence justification for combining the four channels. If the four logical error channels are positively correlated or simply additive, the quoted volume can underestimate the volume needed for the total logical error probability to reach 10^-12. Please state the approximation being made and estimate its effect on the reported volumes.
- [Section 2.9] All simulations are on a torus, and the section explicitly says the authors have not verified the planar X1Z1 code, while the planar X1Y1Z1 code was shown in Ref. [GNM22] to perform essentially as well. Because the motivation is hardware and lattice surgery, the transfer of the torus results to planar surface-code blocks is load-bearing for the practical conclusions. Please either add planar boundary simulations for the key comparisons (at least X1Z1 and X1Y1Z1 under the relevant noise models and decoders) or explicitly restrict the central claims to the torus case.
minor comments (5)
- [Figure 10 caption] The caption contains the typo 'mesurements'; it should be 'measurements'.
- [Sections 3.1 and 3.3] There are typos 'teraqoup' and 'necesarrilly' in the text; please correct them.
- [Figure 15 caption] The caption and axis labels use the placeholder '10□12' instead of typeset superscripts; please fix the rendering.
- [Section 5.2] The text refers to '(nM,mE,h r)-block' where the second dimension should be nE, not mE; please correct the typo.
- [Key point 1.4 and Figure 3 caption] Key point 1.4 says repeating measurements is not worthwhile under SI noise, while the Figure 3 caption notes a possible improvement for X1Y1Z1 in SI noise; please reconcile the wording to avoid an apparent contradiction.
Circularity Check
No circularity: the teraquop-volume ranking is an extrapolated simulation output, not a fitted target or a self-citation chain.
full rationale
The paper's central results (Key points 4.5 and 5.7) are numerical teraquop volumes obtained by simulating memory and stability experiments with Stim, PyMatching, and BeliefMatching, then extrapolating logical error rates to 10^-12 (Section 3.1, Appendix A). The teraquop volume (Definition 3.1) is defined independently of any code ranking: it is the minimal spacetime volume at which the logical error probability is 10^-12, and the reported volumes follow from the simulations rather than being imposed as fit targets. The distance-dependent linear fits in log(p_L) versus d are standard threshold extrapolations; the extrapolation quality is explicitly checked for MWPM in Figure 16, and its limitations are acknowledged in Section 3.3 (equal space/time weighting, schedule and patch effects) and Section 2.9 (planar X1Z1 performance not verified). Prior-work citations ([KdlFT+24] for DCCC definitions, [HB21] for SIGF, [GNFB21]/[GNM22] for noise models and the footprint metric) supply constructions and benchmarks but do not determine the measured error rates or the resulting ranking; the belief-matching decoder is independently implemented in Refs. [HBK+23]/[Hig23]. No equation in the paper defines a prediction in terms of the fitted quantities, and no load-bearing claim rests on a self-citation chain. The stated extrapolation and torus-to-planar caveats are correctness risks, not circularity.
Assumptions & free parameters
free parameters (3)
- Total physical error rate, phenomenological noise =
10^-3
- Circuit-level physical error rates (SD, SI, EM) =
5e-4, 2.5e-4, 2.5e-3
- Bias grid (eta_m, eta_Z) =
{1,4,8,16} x {1,4,8,16}
assumptions (5)
- standard math Stabilizer formalism and detector error model definitions from prior literature are correct for DCCCs.
- domain assumption The three noise models (phenomenological, SI, SD, EM) accurately model relevant hardware with measurement-biased noise.
- domain assumption Logical error rates decay exponentially in the tailored distance and linearly in other dimensions, so independent optimization and line extrapolation to 10^-12 are valid.
- domain assumption Toric simulations are a valid proxy for planar surface-code blocks.
- domain assumption Timelike E and M errors occur roughly equally often in lattice surgery, justifying h = (hE+hM)/2.
Cite this review
Pith. "Pith review of Dynamical codes for hardware with noisy readouts." pith.science (2026). https://pith.science/paper/3LYNKDEO
@misc{pith2026250507658,
author = {Pith},
title = {Pith review of: Dynamical codes for hardware with noisy readouts},
year = {2026},
howpublished = {\url{https://pith.science/paper/3LYNKDEO}},
note = {Machine review of arXiv:2505.07658}
}
abstract
Dynamical stabilizer codes may offer a practical route to large-scale quantum computation. Such codes are defined by a schedule of error-detecting measurements, which allows for flexibility in their construction. In this work, we ask how best to optimise the measurement schedule of dynamically condensed colour codes in various limits of noise bias. We take a particular focus on the setting where measurements introduce more noise than unitary and idling operations - a noise model relevant to some hardware proposals. For measurement-biased noise models, we improve code performance by strategically repeating measurements within the schedule. For unbiased or $Z$-biased noise models, we find repeating measurements offers little improvement - somewhat contrary to our expectations - and investigate why this is. To perform this analysis, we generalise a metric called the teraquop footprint to the teraquop volume. This is the product of the number of qubits and number of rounds of measurements required such that the probability of a spacelike or timelike logical error occurring is less than $10^{-12}$. In most cases, we find differences in performance are primarily due to the number of rounds of measurements required, rather than the number of qubits - emphasising the importance of using the teraquop volume in the analysis. Additionally, our results provide another example of the importance of making use of correlated errors when decoding, in that using belief matching rather than minimum-weight perfect matching can turn a worst-performing code under a given noise model into a best-performing code.
Figures
Figures from the paper (27 more)
Reference graph
Works this paper leans on
-
[2]
[BTHG25] Noah Berthusen, Shi J. S. Tan, Eric Huang, and Daniel Gottesman. Adap- tive syndrome extraction.arXiv preprint arXiv:2502.14835,
-
[5]
[DTTBE24] Peter-Jan H. S. Derks, Alex Townsend-Teague, Ansgar G. Burchards, and Jens Eisert. Designing fault-tolerant circuits using detector error models. arXiv preprint arXiv:2407.13826,
-
[7]
Less bacon more threshold.arXiv preprint arXiv:2305.12046,
[GB23] Craig Gidney and Dave Bacon. Less bacon more threshold.arXiv preprint arXiv:2305.12046,
-
[9]
Stim: Command line usage documentation
[Gid] Craig Gidney. Stim: Command line usage documentation. https://github.com/quantumlib/Stim/blob/main/doc/usage_ command_line.md. Accessed: 2025-01-21. [Gid21] Craig Gidney. Stim: a fast stabilizer circuit simulator. Quantum, 5:497,
work page 2025
-
[10]
Accessed: 2025-01-21. [Gid23a] Craig Gidney. Crumble - (PROTOTYPE) point-and-click 2D QEC circuit builder. https://algassert.com/crumble,
work page 2025
-
[11]
Surviving as a quantum computer in a classical world - 2024 draft,
[Got24] Daniel Gottesman. Surviving as a quantum computer in a classical world - 2024 draft,
work page 2024
-
[12]
Using detector likelihood for benchmarking quantum error correction
[HHW24] Ian Hesner, Bence Hetényi, and James R Wootton. Using detector likelihood for benchmarking quantum error correction. arXiv preprint arXiv:2408.02082,
-
[13]
Improved quantum circuits for elliptic curve discrete log- arithms
[HJN+20] Thomas Häner, Samuel Jaques, Michael Naehrig, Martin Roetteler, and Mathias Soeken. Improved quantum circuits for elliptic curve discrete log- arithms. In Post-Quantum Cryptography: 11th International Conference, PQCrypto 2020, Paris, France, April 15–17, 2020, Proceedings 11, pages 425–444. Springer,
work page 2020
Show all 22 references
-
[15]
Efficient color code de- coders in d ≥ 2 dimensions from toric code decoders
46 [KD19] Aleksander Kubica and Nicolas Delfosse. Efficient color code de- coders in d ≥ 2 dimensions from toric code decoders. arXiv preprint arXiv:1905.07393,
1905 arXiv
-
[17]
Floquetify- ing stabiliser codes with distance-preserving rewrites
[RPK24] Benjamin Rodatz, Boldizsár Poór, and Aleks Kissinger. Floquetify- ing stabiliser codes with distance-preserving rewrites. arXiv preprint arXiv:2410.17240,
-
[18]
Moylett, and Coral M
[SJG+25] Evan Sutcliffe, Bhargavi Jonnadula, Claire Le Gall, Alexandra E. Moylett, and Coral M. Westoby. Distributed quantum error correction based on hyperbolic floquet codes.arXiv preprint arXiv:2501.14029,
-
[19]
Tailoring dynamical codes for 47 biased noise: The X 3Z3 Floquet code
[SM24] Fnu Setiawan and Campbell McLauchlan. Tailoring dynamical codes for 47 biased noise: The X 3Z3 Floquet code. arXiv preprint arXiv:2411.04974,
-
[20]
ZX-calculus for the working quantum computer scientist
[vdW20] John van de Wetering. ZX-calculus for the working quantum computer scientist. arXiv preprint arXiv:2012.13966,
2012 arXiv
-
[22]
Pauli web of the|y⟩ state surface code injection
[WZ25] Kwok Ho Wan and Zhenghao Zhong. Pauli web of the|y⟩ state surface code injection. arXiv preprint arXiv:2501.15566,
-
[1965]
War- ren, Jonathan Gross, et al
[EMS+24] Alec Eickbusch, Matt McEwen, Volodymyr Sivak, Alexandre Bourassa, Juan Atalaya, Jahan Claes, Dvir Kafri, Craig Gidney, Christopher W. War- ren, Jonathan Gross, et al. Demonstrating dynamic surface codes.arXiv preprint arXiv:2412.14360,
-
[2019]
Quantum graphical models and belief prop- agation
[LP08] Matt Leifer and David Poulin. Quantum graphical models and belief prop- agation. Ann. Phys., 323:1899–1946,
1946
-
[2020]
[Woo22] James R. Wootton. Measurements of Floquet code plaquette stabilizers. arXiv preprint arXiv:2210.13154,
-
[2021]
Floquet codes without parent subsystem codes
[DTB22] Margarita Davydova, Nathanan Tantivasadakarn, and Shankar Balasubra- manian. Floquet codes without parent subsystem codes. arXiv preprint arXiv:2210.02468,
-
[2022]
Quantum processing units
[Com] Rigetti Computing. Quantum processing units. https://qcs.rigetti. com/qpus. Accessed: 2025-05-09. [CSB+24] Laura Caune, Luka Skoric, Nick S. Blunt, Archibald Ruban, Jimmy Mc- Daniel, Joseph A. Valery, Andrew D. Patterson, Alexander V. Gramolin, Joonas Majaniemi, Kenton M...
2025 arXiv
-
[2023]
How to factor 2048 bit rsa integers in 8 hours using 20 million noisy qubits.Quantum, 5:433,
[GE21] Craig Gidney and Martin Ekerå. How to factor 2048 bit rsa integers in 8 hours using 20 million noisy qubits.Quantum, 5:433,
-
[2024]
[AMC+24] Hany Ali, Jorge Marques, Ophelia Crawford, Joonas Majaniemi, Marc Serra-Peralta, DavidByfield, BorisVarbanov, BarbaraM.Terhal, Leonardo DiCarlo, and Earl T
Ac- cessed: 2025-03-13. [AMC+24] Hany Ali, Jorge Marques, Ophelia Crawford, Joonas Majaniemi, Marc Serra-Peralta, DavidByfield, BorisVarbanov, BarbaraM.Terhal, Leonardo DiCarlo, and Earl T. Campbell. Reducing the error rate of a supercon- ducting logical qubit using analog rea...
2025
-
[2025]
[KB25] Gilad Kishony and Erez Berg
Accessed: 2025-05-09. [KB25] Gilad Kishony and Erez Berg. Increasing the distance of topological codes with time vortex defects.arXiv preprint arXiv:2502.12236,
2025
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.