Pith. sign in

REVIEW 3 major objections 5 minor 36 references

Decoder Dependence in Surface-Code Threshold Estimation under Digitized Hybrid Continuous-Variable and Discrete Noise

T0 review · 3 major / 5 minor · reviewed 2026-07-15 · grok-4.5

Pith's one-line read Surface-code threshold numbers are outputs of a decoder-and-estimator pipeline, not fixed properties of the code and noise alone.

desk verdict Careful matched-decoder study that makes a real methodological point about threshold auditing under digitized hybrid noise, with hybrid numbers that stay estimator- and fallback-conditional. read the letter →

arxiv 2603.06730 v4 pith:F7M33S5A submitted 2026-03-06 quant-ph

classification quant-ph
keywords surfacecodethresholdestimationdecoderdependencematchingUnion-Findhybridcontinuous-variablenoiselogicalerrorratefinite-sizescaling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that when you estimate a surface-code threshold, the number you get depends on which decoder you use and how you locate the crossing in finite data. The authors run matching-style decoding, Union-Find, and a light neural reweighting of matching inside one shared workflow, first under ordinary Pauli noise and then under a hybrid model that turns continuous Gaussian displacements into discrete faults. Matching consistently beats Union-Find under Pauli noise and yields a crossing near physical error rate 0.05; under hybrid noise the same backends give interior crossings that split by distance pair and stay sensitive to the grid and estimator. At larger distance the matching path falls back to a greedy rule for many shots, so the reported logical-error curves mix exact and approximate behavior. The practical claim is that auditable threshold work must publish decoder choice, sweep resolution, and backend fallback rates alongside the scalar number.

What carries the argument

A controlled LiDMaS+ comparison protocol: identical seeds, distance sets, and noise grids across backends, with hybrid Gaussian noise first mapped to parity-encoded Pauli faults (no continuous soft stream), and with hybrid crossing localization that drops the initial exact-zero logical-error plateau before interpolating sign changes or reporting minimum-separation proxies.

What would settle it

Rerun the same hybrid dense window and distance-9 points with a production minimum-weight matching backend that never falls back to the greedy path, and check whether the (d=5,7) interior crossing stays near 0.33, moves toward the (d=3,5) value, or disappears once fallback rates are near zero.

Watch

Extended reading notes

Core claim

Under a single matched workflow, decoder and estimator choice change surface-code threshold summaries for both Pauli-reference noise and digitized hybrid continuous-variable noise. Matching-style decoding outperforms Union-Find on Pauli sweeps and returns a crossing median pc of 0.0531 with a consistent collapse fit near 0.052; hybrid dense-window interior crossings for matching are about 0.47 for distances 3 and 5 and about 0.33 for distances 5 and 7, with the latter low-error estimate remaining estimator-sensitive, while high fallback rates at distance 9 limit how far the hybrid numbers can be pushed.

Load-bearing premise

That a matching backend which solves only small defect sets exactly and falls back to a greedy rule, plus a hybrid model that feeds only digitized Pauli events without soft information, is still a fair enough stand-in for production matching and for GKP-style continuous noise that decoder-family effects can be read off the reported curves.

Editorial extensions

If this is right

  • Published surface-code thresholds should name the decoder, estimator, grid resolution, and fallback or decoder-failure rates, not only a scalar pc or σc.
  • Union-Find can preserve qualitative distance reversal under hybrid noise while still inflating logical error at larger distance and moderate-to-high noise relative to matching.
  • Hybrid crossings extracted after dropping exact-zero plateaus remain pair-dependent and low-error estimates stay grid-sensitive, so they should not be treated as a single converged critical point.
  • Lightweight learned reweighting of matching edges can modestly lower sampled mean logical error without becoming an end-to-end neural decoder.
  • High matching-fallback rates at larger distance make those logical-error curves partly measures of the fallback policy, not of the exact matching objective alone.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If hybrid architectures keep digitizing continuous noise before decoding, threshold claims will need separate analog-aware and digitized pipelines rather than one shared discrete decoder stack.
  • Adaptive trial budgets concentrated where distance curves nearly touch would likely shrink the gap between coarse-grid proxies and interior hybrid crossings more than simply raising total shots.
  • Once production matching removes greedy fallback, remaining decoder gaps between matching and Union-Find under hybrid noise would cleanly isolate approximation quality from implementation artifacts.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript argues that surface-code threshold estimates are outputs of an inference pipeline and therefore depend on decoder backend, estimator, sweep resolution, and statistical budget. Within a single LiDMaS+ harness with matched seeds, grids, and trial counts, the authors compare a matching-style backend (exact small-instance solver with greedy fallback), Union-Find, and a lightweight neural-guided reweighter under a Pauli-reference channel and a digitized hybrid continuous-variable/discrete channel that maps Gaussian displacements to parity-mapped Pauli faults without soft information. In the Pauli-reference mode they report a matching-style crossing median pc=0.0531 [0.0415,0.0572] and collapse fit pc=0.052 (ν=1.35), while Union-Find yields no stable crossing. In a dense hybrid transition window they report interior matching-style crossings σc=0.4707 (d=3,5) and σc=0.3275 (d=5,7), the latter flagged as low-LER and estimator-sensitive; a d=9 extension shows high matching-fallback rates (up to 0.747) and larger Union-Find LER; a d=5 guidance sweep shows a modest mean-LER reduction under full reweighting. The methodological conclusion is that estimator resolution and backend fallback diagnostics belong in an auditable decoder comparison.

Significance. If the results hold under the stated protocol, the paper makes a useful methodological contribution: it treats threshold estimation as a controlled, reproducible pipeline comparison rather than as a decoder-free material constant, and it documents fallback and decoder-failure diagnostics alongside LER. The matched-seed design, Wilson intervals, explicit exclusion of exact-zero plateaus before crossing localization, and joint reporting of crossing versus collapse summaries for the Pauli control are strengths that support auditability. The absolute Pauli pc≈0.05 is well below standard MWPM bit-flip thresholds, and the hybrid model is deliberately digitized without soft information, so the work does not claim production decoder performance or a full GKP-style analog decoder study. Within that scope, the controlled comparison and the insistence on reporting estimator and fallback metadata are of practical value to groups that publish finite-distance threshold summaries under nonstandard noise maps.

major comments (3)
  1. [Table I / §II hybrid mode] The hybrid noise map is central to half the study (Table I; hybrid mode throughout §§III–IV) but is specified only as “Gaussian→parity-mapped Pauli” with no soft stream. An explicit, self-contained digitization rule (e.g., the quadrature-to-Pauli binning or parity map, including any GKP-style modular reduction and the precise relation between σ and the effective Pauli rate) is needed so that the dense-window σc values and multi-distance reversals can be reproduced without the full LiDMaS codebase. Without that equation, the hybrid crossings remain protocol-internal numbers rather than independently checkable results.
  2. [§II matching description; Table VI; Table VII] The matching-style backend is exact only below a configured small-instance limit and uses a greedy fallback above it; the paper correctly reports a maximum fallback rate of 0.747 at d=9 (Table VII) and states that results are not production Blossom benchmarks. Fallback rates are not tabulated for the dense d=3,5,7 transition window or the coarse multi-distance hybrid sweeps that supply the main decoder-ordering and σc claims (Fig. 3, Table VI, Table IV). Because high-σ, larger-d points are exactly where defect sets grow, those LER curves and the (d=5,7) crossing may mix exact matching with the greedy policy. Reporting fallback rate versus (d,σ) for every hybrid curve used in crossing localization is load-bearing for interpreting “matching-style vs Union-Find” as a decoder-family comparison rather than a fallback-policy comparison.
  3. [Table V / §IV.C Pauli threshold] The Pauli-reference matching-style crossing median pc=0.0531 and collapse pc=0.052 (Table V, Fig. 5) are substantially below the ~10% independent bit-flip MWPM thresholds standard in the literature for the surface code. The manuscript notes that the backend is not production Blossom, but it does not discuss this absolute discrepancy as a control-check on the protocol. A short quantitative discussion—attributing the reduction to the greedy fallback, the restricted noise branch (X→Z only), finite-distance effects, or another identified cause—would strengthen the claim that the Pauli mode is a “stable comparison point” rather than an anomalous internal control.
minor comments (5)
  1. [Figs. 2–5] Several figure panels in the compiled text appear with corrupted axis labels or placeholder glyphs (e.g., Figs. 2–5 in the source dump). Ensure final production figures have legible axis labels, legends, and Wilson-band captions.
  2. [Table IX] Table IX compares heterogeneous published thresholds and metrics; the caption already warns against one-to-one benchmarking, but a column for noise model and decoder family would make the non-comparability more transparent.
  3. [§II neural-guided matching; Table VIII] The neural-guided model is described as a linear reweighter trained simulator-in-the-loop (Alg. 4; Eqs. 7–8 in the appendix), but training set size, regularization, and whether the sensitivity sweep uses a held-out σ grid are not stated. A brief note would clarify that the 0.1773→0.1663 mean-LER change is an internal sensitivity result, not a learned-decoder benchmark.
  4. [Eq. (12); Table VI] Eq. (12) linearizes the crossing between adjacent grid points after plateau exclusion. State explicitly whether the reported σc values use the midpoints of the 0.01 grid or the continuous interpolant, and give the two bracketing (σ, ΔLER) pairs for the (d=5,7) estimate so readers can judge low-LER sensitivity.
  5. [§III–V headings] Minor typography: “EXPERIMENT AL DESIGN”, “RESUL TS”, and “V alidity” spacing artifacts in headings should be cleaned for the camera-ready version.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: threshold summaries are Monte-Carlo LER outputs under fixed protocols, not quantities forced by definition or self-citation.

full rationale

The paper’s load-bearing chain is experimental, not definitional. Physical noise is sampled, syndromes extracted, corrections produced by concrete backends (matching-style with documented greedy fallback, Union-Find, neural-guided reweighting), and LER is the empirical frequency F/T (Eq. 9) with Wilson intervals (Eq. 10). Pairwise crossings use the standard linearized sign-change interpolant (Eqs. 11–12) or a min-separation proxy when no sign change exists (Eq. 13); the Pauli collapse uses the standard finite-size ansatz (Eq. 14). None of these estimators is defined in terms of the reported pc or σc values, so the numerical thresholds are not forced by construction. Neural guidance is an internal λ-sensitivity sweep (λ=0,0.5,1), not a first-principles prediction of an independent observable. The self-citation to LiDMaS (arXiv:2601.16244) identifies the shared simulation harness; it does not supply a prior numerical threshold that this work re-derives. Decoder ordering and hybrid estimator-sensitivity claims are therefore independent Monte-Carlo content within the stated protocol, not circular reductions.

Assumptions & free parameters 5 free parameters · 6 assumptions · 3 invented entities

The central claim rests on standard surface-code and decoding theory plus a set of implementation and estimator choices that define the LiDMaS+ protocol. Free parameters are mostly experimental-design knobs (grids, trials, guidance strength, small-instance cutoff) rather than physics constants fitted to data. Invented entities are the software workflow and the specific digitized hybrid mode; they have no independent experimental handle outside this simulation stack.

free parameters (5)
  • matching small-instance exact-solver cutoff / greedy fallback threshold
    Unstated numeric limit above which the matching-style backend switches from exact subset DP to greedy; directly affects LER and the reported fallback rates up to 0.747 at d=9.
  • neural guidance strength λ and reweighting clip
    λ∈{0,0.5,1} and the clip on αij are chosen by the authors; full λ=1 produces the reported mean-LER drop from 0.1773 to 0.1663.
  • dense hybrid transition window σ∈[0.30,0.50] step 0.01 and 3000 trials/point
    Hand-chosen window and budget that define which interior crossings are visible after plateau exclusion.
  • collapse search region and bootstrap count (100) for Pauli FSS
    Union-Find collapse hits the lower boundary pc=0.040; the search bounds and bootstrap sample size affect the reported interval and cost.
  • base seed s0=1337 and deterministic seed map g
    Fixes the entire Monte-Carlo realization; different seeds would shift finite-sample crossings.
assumptions (6)
  • domain assumption Standard surface-code syndrome extraction and logical-error definition under independent X (or Z) faults
    Used throughout §II–IV as the object whose LER is measured; taken from Kitaev and subsequent literature.
  • domain assumption Finite-size scaling collapse ansatz pL(d,p)≈f((p−pc)d^{1/ν})
    Eq. (14); standard QEC practice, not re-derived.
  • standard math Wilson-score confidence intervals for binomial LER
    Eqs. (9)–(10) and proof sketch in appendix; classical statistics.
  • ad hoc to paper Linearized crossing estimator when ΔLER changes sign between adjacent grid points; min-|ΔLER| proxy otherwise; exclude exact-zero plateaus
    Eqs. (11)–(13) and hybrid protocol text; the plateau exclusion is a paper-specific estimator rule that changes which σc values are reported.
  • ad hoc to paper Digitized hybrid map: continuous Gaussian displacement → parity-mapped Pauli faults with no soft information to the decoder
    Table I and hybrid mode definition; motivated by GKP literature but the concrete map is an implementation choice of this workflow.
  • domain assumption Union-Find and matching (pair/boundary) objectives as stated in Algorithms 2–3
    Standard decoder families; the paper’s matching is not claimed to be production Blossom.
invented entities (3)
  • LiDMaS+ workflow / harness
    purpose: Common syndrome-extraction, decoder, correction, aggregation, and reporting path that makes decoder comparison controlled.
    Authors’ software stack; all results are conditional on this implementation. No independent public verification cited.
  • Digitized hybrid CV/discrete operational mode (parity-mapped Pauli only)
    purpose: Test decoder dependence after continuous noise is mapped to discrete decoder-facing events without soft info.
    Defined in Table I; motivated by GKP but the exact digitization and lack of soft stream are paper-specific.
  • Lightweight linear neural-guided matching reweighter trained simulator-in-the-loop
    purpose: Sensitivity analysis of learned edge reweighting (λ) on hybrid LER.
    Internal model; not a production learned decoder; results are sensitivity only.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Decoder Dependence in Surface-Code Threshold Estimation under Digitized Hybrid Continuous-Variable and Discrete Noise." pith.science (2026). https://pith.science/paper/F7M33S5A

@misc{pith2026260306730,
  author       = {Pith},
  title        = {Pith review of: Decoder Dependence in Surface-Code Threshold Estimation under Digitized Hybrid Continuous-Variable and Discrete Noise},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F7M33S5A}},
  note         = {Machine review of arXiv:2603.06730}
}
abstract

Surface-code threshold estimates depend on the inference pipeline, including decoder and estimator choices. We compare decoders within a single LiDMaS+ workflow under Pauli-reference and digitized hybrid continuous-variable/discrete sweeps. In the Pauli-reference mode, the matching-style backend outperforms Union-Find and yields crossing median $p_c=0.0531$ (bootstrap interval $[0.0415,0.0572]$) and collapse fit $p_c=0.052$ ($\nu=1.35$). For the hybrid mode, a dense transition-window sweep at $d=3,5,7$ uses $\sigma\in[0.30,0.50]$ with step $0.01$ and $3000$ trials per point. After the initial exact-zero plateau is excluded from crossing localization, the matching-style backend gives interior crossing estimates $\sigma_c=0.4707$ for $(d=3,5)$ and $\sigma_c=0.3275$ for $(d=5,7)$; the latter lies in a low-LER region and remains estimator-sensitive. A targeted $d=9$ extension shows larger Union-Find LER at moderate-to-high $\sigma$ and matching-fallback rates up to $0.747$ at $\sigma=0.50$. In a $d=5$ neural-guidance sensitivity sweep, full learned reweighting reduces the sampled mean LER from $0.1773$ to $0.1663$ over $\sigma\in[0.35,0.55]$. These results show that estimator resolution and backend fallback diagnostics are part of an auditable decoder comparison.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

36 extracted references · 1 linked inside Pith

  1. [1]

    Starting from the estimator in Eq

    Logical-Error and Threshold Estimators For code distance d, sweep parameter θ (either p or σ), andTtrials withFlogical failures, �p� (d, θ) = F T .(9) The implementation reports Wilson-score confidence inter- vals. Starting from the estimator in Eq. (9), with z = 1.96 and�p� =F/T, define µ= �p� +z 2/(2T) 1 +z 2/T , h= z 1 +z 2/T � �p� (1��p � ) T + z2 4T ...

  2. [2]

    Consistency Proof Sketches ��������� �������� ������ �� ��� (10) �� ��� ������ ����� �������� ��� � ��������� ���������� ���� ��������� ���������z� ������Let ˆp= �p� = F/T from Eq. (9). The score-test acceptance region for null proportionpis (ˆp�p) 2 p(1�p)/T �z 2. Rearrangement gives a quadratic inequality inp: (T+z 2)p2 �(2Tˆp+z 2)p+Tˆp 2 �0. The soluti...

  3. [3]

    A deterministic seed map s(d, θ, t) =g(s 0, d, θ, t) (15) guarantees identical pseudo-random streams for repeated executions under the same configuration

    Deterministic Reproducibility Parameterization Each run is indexed by ( d, θ, t) with base seed s0. A deterministic seed map s(d, θ, t) =g(s 0, d, θ, t) (15) guarantees identical pseudo-random streams for repeated executions under the same configuration. Together with fixed sweep sets �=�3,5,7�, �=�0.04,0.05, . . . ,0.12�, Σ =�0.05,0.10, . . . ,0.60�, Σre...

  4. [4]

    A. Y. Kitaev, Russian Mathematical Surveys52, 1191 (1997)

  5. [5]

    Gottesman, A

    D. Gottesman, A. Kitaev, and J. Preskill, Physical Review A64, 012310 (2001)

  6. [6]

    Tzitrin, T

    I. Tzitrin, T. Matsuura, R. N. Alexander, G. Dauphinais, J. E. Bourassa, K. K. Sabapathy, N. C. Menicucci, and I. Dhand, PRX Quantum2, 040353 (2021)

  7. [7]

    D. D. K. Wayo, Lidmas: Architecture-level modeling of fault-tolerant magic-state injection in gkp photonic qubits (2026), arXiv:2601.16244 [quant-ph]

  8. [8]

    Higgott and C

    O. Higgott and C. Gidney, Quantum9, 1600 (2025)

Show all 36 references
  1. [9]

    Delfosse and N

    N. Delfosse and N. H. Nickerson, Quantum5, 595 (2021), originally circulated as arXiv:1709.06218 (2017)

  2. [10]

    J. P. Bonilla Ataides, D. K. Tuckett, S. D. Bartlett, S. T. Flammia, and B. J. Brown, Nature Communications12, 2172 (2021)

  3. [11]

    A. S. Darmawan, B. J. Brown, A. L. Grimsmo, D. K. Tuckett, and S. Puri, PRX Quantum2, 030345 (2021)

  4. [12]

    Higgott, T

    O. Higgott, T. C. Bohdanowicz, A. Kubica, S. T. Flammia, and E. T. Campbell, Physical Review X13, 031007 (2023)

  5. [13]

    Skoric, D

    L. Skoric, D. E. Browne, K. M. Barnes, N. I. Gillespie, and E. T. Campbell, Nature Communications14, 7040 (2023)

  6. [14]

    X. Tan, F. Zhang, R. Chao, Y. Shi, and J. Chen, PRX Quantum4, 040344 (2023)

  7. [15]

    Q. Xu, N. Mannucci, A. Seif, A. Kubica, S. T. Flammia, and L. Jiang, Physical Review Research5, 013035 (2023)

  8. [16]

    A. Dua, A. Kubica, L. Jiang, S. T. Flammia, and M. J. Gullans, PRX Quantum5, 010347 (2024)

  9. [17]

    Kobayashi and G

    R. Kobayashi and G. Zhu, PRX Quantum5, 020360 (2024)

  10. [18]

    Sahay, Y

    K. Sahay, Y. Lin, S. Huang, K. R. Brown, and S. Puri, PRX Quantum6, 020326 (2025)

  11. [19]

    Y. Kang, J. Lee, J. Ha, and J. Heo, Quantum Information Processing23, 190 (2024)

  12. [20]

    J. Ha, Y. Kang, J. Lee, and J. Heo, Quantum Information Processing24, 164 (2025)

  13. [21]

    Y. Zhao, Y. Liu, C. Zhang, and et al., Physical Review Letters129, 030501 (2022)

  14. [22]

    Krinner, N

    S. Krinner, N. Lacroix, A. Remm, and et al., Nature605, 669 (2022)

  15. [23]

    Takeda, A

    K. Takeda, A. Noiri, T. Nakajima, T. Kobayashi, and S. Tarucha, Nature608, 682 (2022)

  16. [24]

    Z. Ni, S. Li, X. Deng, and et al., Nature616, 56 (2023)

  17. [25]

    Google Quantum AI, Nature614, 676 (2023)

  18. [26]

    R´ eglade, A

    U. R´ eglade, A. Bocquet, R. Gautier, and et al., Nature 629, 778 (2024)

  19. [27]

    A. Z. Ding, B. L. Brock, A. Eickbusch, and et al., Nature Communications16, 5279 (2025)

  20. [28]

    Google Quantum AI and Collaborators, Nature638, 920 (2025)

  21. [29]

    S. C. Smith, B. J. Brown, and S. D. Bartlett, Communi- cations Physics7, 386 (2024)

  22. [30]

    Bravyi, A

    S. Bravyi, A. W. Cross, J. M. Gambetta, D. Maslov, P. Rall, and T. J. Yoder, Nature627, 778 (2024)

  23. [31]

    Q. Xu, G. Zheng, Y.-X. Wang, P. Zoller, A. A. Clerk, and L. Jiang, npj Quantum Information9, 78 (2023)

  24. [32]

    D. S. Wang, A. G. Fowler, and L. C. L. Hollenberg, Phys- ical Review A83, 020302 (2011)

  25. [33]

    A. G. Fowler, Physical Review Letters109, 180502 (2012)

  26. [34]

    D. K. Tuckett, S. D. Bartlett, and S. T. Flammia, Physical Review Letters120, 050505 (2018)

  27. [35]

    D. K. Tuckett, S. D. Bartlett, S. T. Flammia, and B. J. Brown, Physical Review Letters124, 130501 (2020)

  28. [36]

    J. Lee, J. Park, and J. Heo, Quantum Information Pro- cessing20, 231 (2021)

Pith tools

Reviewed July 15, 2026 · model on record in the stance chip above.