{"id":"b8bb88b3-ca0d-44f2-a16f-24b2ab8534c9","arxiv_id":"2607.19447","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A four-branch syndrome-to-decoder validation shows repetition and Steane CSS codes preserve intended syndromes on IBM hardware, while a distance-five surface-code Z-check layer is overwhelmed by routing noise and a PennyLane GKP proxy retains partial target localization.","lead":"This paper runs four quantum error-correction setups—three on IBM hardware and one simulated GKP model—and checks whether raw syndrome bits can be fed into a standard decoder pipeline without losing code semantics. It is an interface-validation study, not a threshold or new-code result, and reports that small codes keep the intended syndromes while a 56-qubit surface circuit is swamped by hardware noise.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Surface-code branch's mapping of 56 measured bits to 40 data qubits and 16 Z-check ancillas is unverified; with exact localization 0.003–0.108, a bit-order error would be indistinguishable from hardware noise and would invalidate the central interface claim at scale.","rationale":"We considered several concerns: the GKP noise parameters are undisclosed; the zero-residual audit in Table IV is guaranteed by Algorithm 1's closure step; and the surface bit-order mapping is unverified. The mapping is the most load-bearing because it is the only issue that could directly invert the interpretation of the paper's largest hardware result. Repetition and Steane branches are compelling evidence that the parser works for circuits with clear dominant syndromes; they do not constrain a 56-qubit routed circuit where the dominant syndrome is absent. The local reference cannot detect a systematic mapping error because it uses the same parser. Figure 14's 'empirical noise diagnostic' presupposes the mapping. If the mapping is wrong, the surface branch does not show that hardware noise defeats exact localization; it shows the interface failed to parse. The proposed concrete test (inspect creg ordering, or run a noiseless transpiled-circuit injection) would settle this, as it gives an unambiguous ground truth. We therefore keep the reader's CONDITIONAL verdict: the paper's central claim is plausible but depends on verification of the surface mapping. We agree with the reader's weakest_assumption.","tokens_in":12413,"tokens_out":4782,"duration_ms":53280,"concrete_test":"Re-derive the surface mapping from the compiled circuit: after transpilation, extract the measurement creg order from the Qiskit compiled circuit (e.g., `circuit.qasm()` and the creg list) and compare it to the assumed 40-data + 16-ancilla logical order. A second independent check: inject X on a low-weight boundary data qubit (e.g., target 1 or 5) and run the exact transpiled circuit in a noiseless simulator; confirm the dominant 16-bit measurement equals that column of H_surf,Z. If the mapping differs or the syndrome does not match, the surface branch's broad activation is at least partly a parse artifact, and the paper's interpretation of 'defeats exact localization' would need revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section II.B states that the pipeline 'map[s] measured bits into the logical row order of H_F' (Algorithm 1), but for the surface-code branch this mapping is not independently demonstrated. The 56-qubit transpiled circuit (Fig. 4) has backend-ordered classical bits; the parser reorders them into 40 data + 16 Z-ancilla order. If any index is permuted, every syndrome column is shuffled, producing broad activation across all 16 Z checks that would look exactly like the observed hardware background in Fig. 7. The local simulator reference uses the same parser, so it cannot rule out a shared systematic misordering. The paper attributes the broad activation to 'hardware-induced syndrome activation rather than a check-indexing error' (Sec. III.B), but no evidence distinguishes these hypotheses. The replay audit in Table IV cannot help, since Algorithm 1's closure step ('close any residual syndrome diagnosed by H_F \\hat{e}_d != s_r') guarantees zero residuals by construction. Thus the central claim that the interface remains semantically aligned at the 56-qubit scale rests on an unverified assumption. Repetition (Sec. III.A) and Steane (Sec. III.D) branches confirm the mapping only for small circuits, where dominant syndromes are unambiguous; they do not test the surface transpilation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a four-branch syndrome-to-decoder validation study: a five-qubit repetition code, a distance-five rotated-surface-code Z-check layer, the Z-check half of the Steane CSS code as a compact CSS-LDPC benchmark, and a PennyLane-backed digitized-GKP companion study. For each branch, the authors build clean and injected circuits, execute 4096-shot streams on IBM hardware (or sample PennyLane Gaussian-CV records for GKP), parse measured bits into LiDMaS+ requests, replay MWPM, UF, and BP policies, and report correction-localization rates with Wilson confidence intervals. The two smaller hardware families preserve the expected dominant syndromes; the 56-qubit surface run shows broad hardware-induced activation with exact localization 0.003--0.108; the GKP branch demonstrates that analog readouts can be binned into the same binary interface. The paper explicitly positions itself as an auditable interface methodology rather than a threshold demonstration.","tokens_in":12751,"tokens_out":8086,"duration_ms":95641,"significance":"If the central interface claim is sound, the paper offers a useful and reproducible benchmark: the code is released, the pipeline steps are documented with make targets, shot counts are fixed, Wilson intervals are reported, and the surface-code failure mode is described honestly rather than dressed up as a threshold result. The repetition and CSS-LDPC branches are clean positive evidence that the parser/check-matrix mapping can be validated on small circuits. However, the surface branch's claimed semantic alignment rests on an unverified bit-order/check-index mapping, and the 'replay audit' in Table IV is circular because Algorithm 1 closes residual syndromes before recording them. These two issues are load-bearing for the paper's full claim, so the manuscript needs revision before the 56-qubit-scale interface claim can be accepted.","major_comments":[{"comment":"The pipeline explicitly includes a syndrome-closing step ('close any residual syndrome diagnosed by H_F e_d != s_r') before recording policy diagnostics. Table IV's zero residual count for every MWPM/UF/BP row is therefore guaranteed by construction and cannot be used as evidence that the decoder policies correctly close the measured syndromes. This weakens the 'replay audit' language in Section III.E and Section IV. Please report pre-closure residuals, or state plainly that residual zero is a post-closure invariant rather than an independent audit.","section":"Algorithm 1 / Table IV"},{"comment":"For the 56-qubit surface branch, the mapping from the transpiled circuit's 56 measured classical bits to the parser's logical order (40 data + 16 Z-check ancillas) is not independently verified. The local simulator uses the same parser, so a systematic bit-order or check-index permutation would make both the hardware and local heatmaps look column-like, and in the hardware data the broad activation produced by such a permutation would be indistinguishable from the reported 'hardware-induced syndrome activation'. With exact localization only 0.003--0.108, Fig. 7 contains no unambiguous column structure to certify the mapping. The sentence 'consistent with hardware-induced syndrome activation ... rather than a check-indexing error' is an assertion, not a test. Please provide an explicit check of the backend measurement ordering against H_surf,Z, for example per-target per-Z-check expected-","section":"Section III.B, Fig. 7"},{"comment":"The minimum-weight objective in Eq. (4) can have multiple minimizers for the degenerate surface-code syndrome, but the paper does not specify a tie-breaking rule for the MWPM/minimum-weight baseline. Exact localization ('e_hat = {i}') is therefore not uniquely defined without the code's tie-breaking convention; for the surface branch the reported 0.003--0.108 range could depend strongly on that choice. Please state the tie-breaking rule and, ideally, report the rate at which the intended singleton is one of the minimum-weight corrections, not merely whether it is the selected representative.","section":"Eq. (4) / Sections III.B and III.C"}],"minor_comments":[{"comment":"The GKP digitization uses an ad hoc decision window |y_j| <= 0.25 sqrt(pi), a q-shift of 0.56 sqrt(pi), and a Gaussian proxy rather than finite-energy non-Gaussian grid states. The paper is appropriately candid that this is a proxy, but it should explicitly state these are model choices and report sensitivity to the window half-width and noise parameters.","section":"Section II.C"},{"comment":"The title 'Hardware-in-the-Loop ... and Digitized-GKP Codes' may overstate the GKP branch, which is a PennyLane model study rather than a hardware-in-the-loop experiment. Consider adding a qualifier such as 'companion digitized-GKP study' in the title or abstract for accuracy.","section":"Title/Abstract"},{"comment":"The make-based reproducibility protocol is welcome. For a stable benchmark, include a commit hash or versioned release tag, and record the IBM Runtime/backend snapshot date and calibration information so the hardware results can be interpreted and reproduced by others.","section":"Appendix A"},{"comment":"The study-level comparison across code families in Fig. 12 aggregates cases with very different degeneracy and syndrome structure. The comparison is useful as an interface check, but the caption should caution that aggregate localization is not a decoder-quality comparison across code families.","section":"Section IV / Fig. 12"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi — quick take on arXiv:2607.19447. Worth your time if you work on the syndrome-to-decoder boundary in QEC experiments. It contributes a reproducible LiDMaS+ pipeline and runs four code families through it: five-qubit repetition, Steane CSS-LDPC Z-checks, distance-five surface Z-checks, and a digitized-GKP proxy. The clean result is that repetition and Steane branches preserve the expected dominant syndrome and correction on ibm_fez hardware, with Wilson intervals. That's credible evidence the parser, bit ordering, and check indexing work on small circuits. The 56-qubit surface branch is honestly reported as noise-dominated, with exact localization 0.003–0.108.\n\nWhat's new is the combined four-branch study behind a single interface, with code and data shipped. That's a useful infrastructure contribution for auditing syndrome streams before they reach a decoder, and the authors don't oversell it: no threshold claims.\n\nSoft spots: the surface branch's mapping from 56 transpiled classical bits to the 16 Z-check order is load-bearing and never independently verified. With localization that low, a systematic bit-order perm would look identical to the broad activation they attribute to hardware noise. The matched local simulator shares the same parser, so it cannot rule out a shared misordering. The paper calls the run a scaling test, but the scaling test only means something if the mapping is known right. They need a sanity check that's sensitive to ordering — for example, injecting a fault that produces a multi-bit syndrome pattern with a unique signature, or comparing the clean hardware syndrome distribution against a permutation-mismatched reference. Also the local surface baseline itself isn't explained: how does 'local' get 0.72 exact localization? Ideal circuit? Noise model? That should be stated.\n\nTable IV's zero-residual audit is circular: Algorithm 1 closes any residual syndrome before recording, so nonzero residuals are zero by construction. They do disclose the closure step, but the table is easy to over-read as decoder validation rather than interface-closure validation.\n\nThe digitized-GKP branch is off-hardware and under-specified. The text gives decision-window and injected shift sizes, but not the noise parameters (shift noise sigma, half-cell jump rate, flip probability). The appendix has make commands but the actual parameter values are in code. As reported, those localization numbers aren't reproducible. Fixable.\n\nOverall: a careful, honest paper that delivers a reproducible toolchain and solid small-code evidence. The surface mapping and GKP parameter disclosure need to be addressed; then I'd be comfortable citing it as an interface-validation reference. It deserves a serious referee — conditional accept, not reject. I'd bring it to the reading group.","headline":"A useful, honest syndrome-to-decoder validation pipeline with clean small-code evidence; the surface branch's bit-order mapping is unverified and the GKP parameters are undisclosed, but both are fixable.","tokens_in":13237,"tokens_out":3428,"would_cite":true,"duration_ms":35841,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Hardware syndrome records can be parsed and audited against the intended check matrix, with repetition and CSS-LDPC runs preserving the expected corrections and the surface run falling back to target-containing localization.","keywords":["syndrome extraction","quantum error correction","decoder interface","hardware-in-the-loop","surface code","repetition code","CSS-LDPC","GKP readout"],"falsifier":"Run one low-weight surface-code injection with the same transpilation and retrieval, and check whether the most frequent measured syndrome matches the predicted column of the 16-by-40 Z-check matrix. A mismatch with the predicted column—beyond the observed noise background—would show the interface mapping is wrong; the paper reports no dominant-syndrome match for the surface branch, so this remains untested.","tokens_in":12283,"feed_emoji":"⚛️","tokens_out":5782,"duration_ms":56965,"temperature":0.7,"pith_summary":"The paper tries to establish that the boundary between quantum hardware readout and a classical decoder—the place where bit-order, check-indexing, and transpilation errors hide—can be made auditable and executable. It builds three hardware syndrome-extraction circuits (a five-qubit repetition code, a 40-data-qubit distance-five surface-code Z-check layer, and the Steane CSS code as a compact CSS-LDPC benchmark) plus an off-hardware digitized Gottesman-Kitaev-Preskill (GKP) proxy, and replays 4,096-shot syndrome streams through three decoder policies. For repetition and CSS-LDPC, the dominant measured syndrome equals the predicted column of the parity-check matrix for every injected error, and hardware localization stays near 0.82–0.87. The routed 56-qubit surface branch produces broad check activation, so exact localization drops to 0.003–0.108 while target-containing localization remains 0.279–0.642. The paper's conclusion is an auditable syndrome-to-decoder interface, not a threshold claim.","feed_headline":"Repetition and CSS syndromes survive real-hardware decoder audit","feed_subtitle":"Surface code at 56 qubits loses exact localization; digitized GKP readouts enter the same decoding pipeline.","key_machinery":"The carrying mechanism is a replay pipeline that expands per-shot hardware counts into syndrome records, constructs a decoder request from each record plus code metadata, and scores corrections with three policies (minimum-weight perfect matching as the baseline, union-find, and hard-decision belief propagation). The load-bearing identity is the linear check relation s = H e (mod 2): each injected target predicts one column of H, and exact localization is counted when the decoded correction equals that target. The pipeline audits bit order, check order, and decoder dispatch simultaneously.","core_discovery":"For repetition and CSS-LDPC circuits executed on a live superconducting processor, the most common syndrome after injecting a known single-qubit X error is exactly the corresponding column of the intended parity-check matrix, and the minimum-weight decoder returns the injected site as the dominant correction. For the distance-five surface code, the same parser and decoder replay preserve metadata alignment, but the measured syndromes are so broadly activated that exact single-target localization is only 0.003–0.108, which the paper attributes to hardware-induced syndrome activation rather than to a check-mapping error. The digitized-GKP branch shows that analog quadrature readouts, after bei","pith_inferences":["The surface-code branch does not by itself confirm the intended check-matrix mapping: with exact localization at 0.003–0.108, a bit-order or check-index error would be indistinguishable from broad hardware noise.","A dedicated mapping test—running a single low-weight surface injection and comparing the dominant measured syndrome to the predicted 16-bit column—would separate mapping failure from noise; the current data cannot.","The recorded surface streams are a ready benchmark for weighted or calibration-aware decoders, which could raise target-containing localization beyond the unweighted minimum-weight baseline."],"forward_implications":["Future quantum error-correction experiments can use the same request boundary to compare decoders on identical recorded syndrome streams.","Repeated rounds and calibration-aware weights can be added to the same interface without changing the audit structure.","GKP-style bosonic readouts can enter the same binary decoding pipeline after explicit digitization.","The surface-code branch provides a concrete scale at which routing and measurement noise dominate one-round exact localization.","The methodology shifts reported QEC results from threshold claims toward verifiable interface correctness."],"fun_headline_variants":["Decoder audit: repetition and CSS pass, surface code fails exact","Hardware decoder test: repetition and CSS hold, surface code loses","Syndrome-to-decoder check: repetition and CSS OK, surface misfires","Real-chip decoder validation: repetition and CSS pass, surface fuzzy","Decoder interface audit: repetition and CSS survive, surface blur"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The surface-code branch's interface claim depends on the assumption that the 56 measured classical bits from the transpiled circuit map exactly into the parser's assumed order of 40 data qubits and 16 syndrome-measurement ancillas, and a mapping error would look exactly like the broad noise the paper reports.","fun_headline_variants_meta":{"raw":{"variants":["Decoder audit: repetition and CSS pass, surface code fails exact","Hardware decoder test: repetition and CSS hold, surface code loses","Syndrome-to-decoder check: repetition and CSS OK, surface misfires","Real-chip decoder validation: repetition and CSS pass, surface fuzzy","Decoder interface audit: repetition and CSS survive, surface blur"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000214,"raw_usage":{"total_tokens":1305,"prompt_tokens":828,"completion_tokens":477,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":572,"completion_tokens_details":{"reasoning_tokens":388}},"tokens_in":572,"tokens_out":477,"duration_ms":5317,"temperature":1.0,"reasoning_tokens":388,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T13:36:14.357236+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run one low-weight surface-code injection with the same transpilation and retrieval, and check whether the most frequent measured syndrome matches the predicted column of the 16-by-40 Z-check matrix. A mismatch with the predicted column—beyond the observed noise background—would show the interface mapping is wrong; the paper reports no dominant-syndrome match for the surface branch, so this remains untested.","supporting_citations":[],"review_version":1}