{"id":"ce701ffd-3114-4bbb-8e54-4f36d88ff1e4","arxiv_id":"2505.07658","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Repeating measurement rounds in dynamically condensed colour codes substantially reduces teraquop volume under measurement-biased noise, while decoder choice determines whether XZ-only or XYZ schedules win.","lead":"Quantum error correcting codes that are driven by a repeating measurement schedule can be retuned to handle hardware where readout errors are the dominant noise source. The authors show that repeating measurements shrinks the resource cost when measurement noise dominates, and that the best schedule depends on the decoder used.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Teraquop-volume ranking is extrapolated from small, sparsely sampled distances; a bend or slope error at d>16 could flip the claimed ordering.","rationale":"The paper is a careful numerical study with reproducible code, and it honestly discloses the torus limitation. The reader's acceptance rests on the extrapolation being safe. My review finds that the extrapolation is the weakest point because the central ranking claims are quantitative statements about 10^-12 behaviour, and the data do not directly probe that regime for the key decoder/code combinations. The 10^4-shot budget for belief matching is particularly important because Key point 4.5 depends on belief matching; with roughly 1-10 errors at the largest distances, the fitted slope has substantial uncertainty, and the extrapolated crossing can move by more than the reported volume differences. A conditional verdict is appropriate: accept the methodology and qualitative conclusions, but require a large-distance and shot-count verification of the specific orderings before these resource numbers are used for hardware design. If the proposed check confirms the ordering, the ACCEPT verdict stands.","tokens_in":36523,"tokens_out":13275,"duration_ms":140296,"concrete_test":"Run E/M memory and stability experiments for the four representative codes (X1Z1 and X1Y1Z1, each unrepeated and with the best repeated schedule, e.g. X3Z3 and X3Y3Z3) under phenomenological noise with measurement bias eta_m=16, eta_Z=1, using both MWPM and belief matching. Collect at least 10^7 shots per data point for belief matching. Add distances d=20, 24, 32 (and for EM noise add d=12, 16) and refit log(p_L) vs d using only d>=16. If, at the 10^-12 crossing, X1Y1Z1 with belief matching is still the smallest volume and the repeated code is still smaller than the unrepeated code under measurement bias, the central claim holds; if the ordering changes or the fit is no longer linear within error bars, the headline result should be reported as conditional on larger-distance verification.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claims (Key points 4.5 and 5.7) assert specific teraquop-volume orderings with order-of-magnitude differences at error rate 10^-12. These volumes are not measured; they are extrapolations. Section 3.1 fits log(p_L) as a linear function of distance d using data at d = 4,8,12,16 for phenomenological/SD/SI noise and d = 2,4,6,8 for EM noise (Appendix A), then extends the line to 10^-12. The extrapolation assumes (i) the asymptotic exponential regime has already set in by d=16 (d=8 for EM) and (ii) each logical-error channel is exactly linear in the two non-relevant dimensions; Figure 16 verifies this only for two codes under MWPM, not for belief matching or repeated codes. Both assumptions are load-bearing because the reported ranking differences are often only factors of roughly 2-10 in the teraquop height h (Key point 4.6); a modest change in the fitted slope or a curvature term can reorder the codes. The statistical basis is additionally thin for exactly the decoder that produces the headline result: belief-matching points use only 10^4 shots (Appendix A), so at the largest distances the estimated p_L may have relative errors of order 100% or be censored at zero errors. A non-exponential error floor, for example from BP approximation in belief matching on loopy hypergraphs, would make the extrapolated volume infinite and reverse Key point 4.5. The torus-versus-planar gap (Section 2.9) is a separate transfer concern; the extrapolation issue is load-bearing even for the toric claims as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies dynamically condensed colour codes (DCCCs) under noise models with tunable measurement bias and Z bias, and introduces the teraquop volume as a spacetime generalization of the teraquop footprint. Using Stim/Sinter with PyMatching (MWPM) and BeliefMatching, it simulates XaYbZc and XaZb codes on a torus for phenomenological, standard depolarizing, superconducting-inspired, and entangling-measurement noise. It reports that (i) MWPM favors XaZb schedules while belief matching favors XaYbZc schedules, (ii) differences in volume are dominated by the number of measurement rounds rather than qubit count, and (iii) repeating measurements helps mainly under strong measurement bias. The paper also provides analytic distance and hyperedge-likelihood tables and makes the simulation code publicly available.","tokens_in":36867,"tokens_out":7442,"duration_ms":73877,"significance":"If the reported teraquop-volume rankings are correct, the paper gives a concrete, decoder- and hardware-aware rule for choosing DCCC schedules, and the teraquop volume is a useful metric for comparing spacetime overhead. The strengths are the use of a standard, externally benchmarked simulation pipeline (Stim/Sinter/PyMatching/BeliefMatching) with explicit noise models and error bars; the public code for full reproducibility; and the transparent analytic tables (Tables 3, 4, and 6) that separate distance effects from decoder-dependent hyperedge effects. The central caveat is that the headline volumes and orderings are extrapolated from small simulated distances, so the quantitative ranking is less secure than the qualitative trends. The finding that volume differences come predominantly from measurement rounds rather than footprint is an important caution for metric choice in QEC resource estimation.","major_comments":[{"comment":"The headline volumes and rankings are extrapolated, not directly simulated. The fits use distances d=4,8,12,16 (d=2,4,6,8 for EM noise) and assume log p_L is linear in d with linear dependence on the other two dimensions, then extend to 10^-12. Figure 16 validates the linear-in-nM and linear-in-h dependence only for E memory experiments under MWPM for the X1Z1 and X1Y1Z1 codes; no equivalent check is shown for belief matching or for repeated codes. The belief-matching fits use 10^4 shots per point (Appendix A), so at the largest d a run may have zero or very few logical errors, making the fitted slope and intercept carry large relative uncertainty. Since the ranking differences in Key points 4.5 and 5.7 are factors of a few in the teraquop height, a modest change in slope or a bend at d>16 could reorder the codes. Please add data at larger distances or provide a bounded extrapolation analysis, report the number of logical errors per point for the belief-matching fits, and state whether the linear-dependence assumption has been checked for the repeated schedules and for belief matching.","section":"Section 3.1"},{"comment":"The definition requires the probability of any spacelike or timelike logical error to be at most 10^-12, but the estimation procedure minimizes each of nE, nM, hE, and hM separately from four different experiments and then forms the product nE * nM * (hE+hM)/2. The text does not give a union bound or an independence justification for combining the four channels. If the four logical error channels are positively correlated or simply additive, the quoted volume can underestimate the volume needed for the total logical error probability to reach 10^-12. Please state the approximation being made and estimate its effect on the reported volumes.","section":"Definition 3.1"},{"comment":"All simulations are on a torus, and the section explicitly says the authors have not verified the planar X1Z1 code, while the planar X1Y1Z1 code was shown in Ref. [GNM22] to perform essentially as well. Because the motivation is hardware and lattice surgery, the transfer of the torus results to planar surface-code blocks is load-bearing for the practical conclusions. Please either add planar boundary simulations for the key comparisons (at least X1Z1 and X1Y1Z1 under the relevant noise models and decoders) or explicitly restrict the central claims to the torus case.","section":"Section 2.9"}],"minor_comments":[{"comment":"The caption contains the typo 'mesurements'; it should be 'measurements'.","section":"Figure 10 caption"},{"comment":"There are typos 'teraqoup' and 'necesarrilly' in the text; please correct them.","section":"Sections 3.1 and 3.3"},{"comment":"The caption and axis labels use the placeholder '10□12' instead of typeset superscripts; please fix the rendering.","section":"Figure 15 caption"},{"comment":"The text refers to '(nM,mE,h r)-block' where the second dimension should be nE, not mE; please correct the typo.","section":"Section 5.2"},{"comment":"Key point 1.4 says repeating measurements is not worthwhile under SI noise, while the Figure 3 caption notes a possible improvement for X1Y1Z1 in SI noise; please reconcile the wording to avoid an apparent contradiction.","section":"Key point 1.4 and Figure 3 caption"}],"recommendation":"major_revision","confidential_remarks":"The qualitative trends are likely correct and the paper is thorough, but the quantitative ordering rests on extrapolations from small distances, and the statistical basis for the belief-matching fits is thin. I would be willing to accept after the authors add either larger-distance data or a bounded extrapolation analysis for the key claims, and after clarifying the union-bound issue in the teraquop-volume definition. The planar transfer gap should also be addressed or explicitly scoped out."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is a well-executed numerical paper that will be useful to people working on Floquet/dynamical codes, but the headline rankings are extrapolated from small lattices to 10^-12 error rates, so I would take the precise ordering as provisional.\n\nThe genuinely new content is the systematic sweep: the teraquop volume (a natural spacetime extension of teraquop footprint), the application of schedule-induced gauge fixing (repeating measurements) to DCCCs, and the demonstration that under measurement-biased noise repeating helps while under Z-bias or unbiased it doesn't. The decoder story is also clean: all measurement errors in the X1Y1Z1 code are hyperedges, which is bad for MWPM and good for belief matching, and the numerics bear that out. The paper ships code and uses the standard Stim/Sinter/PyMatching/BeliefMatching stack, so the numbers are reproducible in principle. Self-citations to prior constructions are appropriate.\n\nSoft spots: the teraquop volumes are not measured but extrapolated. The fits use d=4,8,12,16 (d=2,4,6,8 for EM noise) and assume exponential decay in distance stays straight to 10^-12. The rankings sometimes differ by factors of 2-10 in block height, so a slope change or curvature term could reorder codes. The stress-test point about belief matching using only 10^4 shots is fair: at the largest distances, that gives maybe 100 expected errors at p_L ~ 10^-2, and at lower rates you can get zero observed errors, which makes line fitting delicate. I don't think this invalidates the qualitative conclusions; the effects are large and consistent. But I wouldn't anchor on the exact volume numbers. Also, all simulations are on a torus; the authors note planar X1Z1 performance is unverified. That's an acknowledged limitation, not a hidden one.\n\nWho is this for? People choosing between DCCC schedules for a specific hardware noise model, and anyone citing resource metrics. It deserves a serious referee; the referee should ask for a sensitivity check on the extrapolation (e.g., fit with and without the largest distance, or report where slopes change), but I would not desk-reject it.\n\nVerdict: accept with revisions.","headline":"Useful, honest numerical study of DCCC schedule tailoring; the teraquop-volume ordering rests on extrapolation and should be read as provisional.","tokens_in":37402,"tokens_out":3106,"would_cite":true,"duration_ms":29726,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Repeating measurements in a dynamical quantum error-correcting code shrinks its spacetime cost when readout noise dominates, but not otherwise.","keywords":["dynamically condensed colour codes","teraquop volume","honeycomb code","belief matching","minimum-weight perfect matching","measurement-biased noise","schedule-induced gauge-fixing","lattice surgery"],"falsifier":"Run the same memory and stability experiments at substantially larger distances, say distance 24 or 32 for phenomenological noise, and check whether logical error rates continue the exponential decay that the line fits assume. A second check is to simulate the X1Z1 code with planar boundaries and compare its logical error rate to the toric result; if planar performance differs markedly, the volume rankings built on toric simulations would need revisiting.","tokens_in":36351,"feed_emoji":"⏱️","tokens_out":5149,"duration_ms":45312,"temperature":0.7,"pith_summary":"This paper asks how to choose the measurement schedule of a dynamically condensed colour code, a family of dynamical stabilizer codes that can serve as the basic block of a lattice-surgery quantum computer, when the hardware's weak point is noisy readout. To answer it, the authors introduce the teraquop volume: the number of qubits times the number of measurement rounds needed to bring the logical error probability below $10^{-12}$. They find that repeating measurements in the schedule improves the volume substantially when the noise is biased toward measurement errors, but has negligible effect for unbiased or Z-biased noise. They also find that the choice of decoder matters as much as the choice of code: belief matching makes the X1Y1Z1 code the best performer while minimum-weight perfect matching makes it the worst.","feed_headline":"Repeating readouts shrinks quantum code volume under measurement noise","feed_subtitle":"Schedule choice and decoder can change spacetime overhead by orders of magnitude.","key_machinery":"The central object is the measurement schedule of a dynamically condensed colour code, specified by how many consecutive rounds of XX, YY, and ZZ edge measurements are performed, giving XaYbZc or XaZb codes. Repeating measurements is schedule-induced gauge-fixing: it creates two-measurement edge detectors that catch measurement errors directly, at the cost of lengthening the face detectors and worsening timelike distances for data-qubit errors. The paper's performance metric is the teraquop volume, the product of qubit count and measurement rounds needed to reach a $10^{-12}$ logical error rate. The comparison tool is the decoding graph, where measurement errors that violate four detectors become hyperedges; belief matching pre-processes these hyperedge probabilities before minimum-weight matching, recovering information that ordinary MWPM throws away.","core_discovery":"The paper's central claim is that the optimal DCCC measurement schedule is not intrinsic to the code but is co-determined by noise bias and decoder. In all three phenomenological noise models studied, the X1Y1Z1 code under belief matching has the smallest teraquop volume, while the same code under MWPM has the largest; the effect grows with measurement bias. Repeating measurements, a form of schedule-induced gauge-fixing that creates two-measurement edge detectors, improves the teraquop volumes of both the XaZb and XaYbZc code families under both decoders when the noise is measurement-biased, but is negligible otherwise, contrary to the authors' initial expectations. At the circuit level, the superconducting-inspired noise bias is not strong enough to make repetition worthwhile, while under entangling-measurement noise the X2Z2 code wins with MWPM. Across most of the parameter sweep, performance differences come primarily from the number of measurement rounds required rather than the number of qubits, which is why the volume metric matters.","pith_inferences":["For a readout-limited device, the practical recipe suggested by these results is a measurement-repeated X1Y1Z1 schedule decoded with belief matching; the paper does not spell this out as a single recommendation, but it follows from the per-bias best-code tables.","The teraquop volume metric could be used to benchmark other spacetime codes against DCCCs, but this paper only compares codes within the DCCC family.","At Z bias beyond 16, asymmetric repetition schedules that spend more time in detectors sensitive to Z errors may begin to pay off; the paper's parameter sweep stops too early to see this.","The reported ranking may be sensitive to the assumption that timelike E and M errors occur equally often in a lattice-surgery computation; if one type dominates, the optimal height differs from the averaged value used here."],"forward_implications":["Lattice-surgery resource estimates that use only the teraquop footprint will miss most of the difference between codes, since the gaps come primarily from the number of measurement rounds.","Hardware with measurement-dominated noise should use schedules with repeated measurements; hardware with unbiased or Z-biased noise gains little from repetition.","Decoder choice should be made jointly with code choice, because belief matching can turn a worst-performing code into the best-performing one under the same noise model.","Under entangling-measurement noise and MWPM, the X2Z2 schedule wins, showing that the optimal schedule also depends on how the multi-qubit measurement is physically implemented."],"supporting_citations":[{"why":"Introduces the teraquop footprint that this paper generalises and supplies the superconducting-inspired noise model.","marker":"[GNFB21]"},{"why":"Defines dynamically condensed colour codes, their boundaries, and the detector patterns used here.","marker":"[KdlFT+24]"},{"why":"Introduces schedule-induced gauge-fixing, the basis for repeating measurements.","marker":"[HB21]"},{"why":"Supplies the belief-matching decoder that changes which code performs best.","marker":"[HBK+23]"},{"why":"Provides the Stim simulator used to construct circuits and sample logical error rates.","marker":"[Gid21]"},{"why":"Benchmarks the planar honeycomb code and supplies the entangling-measurement noise model.","marker":"[GNM22]"},{"why":"Provides the sparse blossom minimum-weight perfect-matching decoder used for MWPM simulations.","marker":"[HG25]"},{"why":"Introduces memory and stability experiments, the methodology used to estimate teraquop dimensions.","marker":"[Gid22a]"}],"fun_headline_variants":["Measurement noise makes repeating readouts pay off in quantum codes","Belief matching can turn worst dynamical code into best","Teraquop volume shows schedule beats qubit count for overhead","Repeating measurements only helps when readouts are noisy","Optimal dynamical code schedule depends on noise bias and decoder"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The teraquop volumes are extrapolated from logical error rates at small code distances down to $10^{-12}$ by assuming exponential decay in one code dimension and linear dependence on the others; if error rates bend at larger sizes, the reported volumes and rankings could change. The simulations also assume a torus, and planar performance for the X1Z1 code has not been verified.","fun_headline_variants_meta":{"raw":{"variants":["Measurement noise makes repeating readouts pay off in quantum codes","Belief matching can turn worst dynamical code into best","Teraquop volume shows schedule beats qubit count for overhead","Repeating measurements only helps when readouts are noisy","Optimal dynamical code schedule depends on noise bias and decoder"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000265,"raw_usage":{"total_tokens":1640,"prompt_tokens":1009,"completion_tokens":631,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":625,"completion_tokens_details":{"reasoning_tokens":550}},"tokens_in":625,"tokens_out":631,"duration_ms":5785,"temperature":1.0,"reasoning_tokens":550,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:11:27.503422+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same memory and stability experiments at substantially larger distances, say distance 24 or 32 for phenomenological noise, and check whether logical error rates continue the exponential decay that the line fits assume. A second check is to simulate the X1Z1 code with planar boundaries and compare its logical error rate to the toric result; if planar performance differs markedly, the volume rankings built on toric simulations would need revisiting.","supporting_citations":[],"review_version":1}