{"id":"45f15dcc-6101-4007-807a-1f4b73619d89","arxiv_id":"2608.02944","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"Sparse error detection in small Iceberg codes reduces systematic errors in simulated Schwinger-model observables under depolarizing noise, with diminishing returns after a few detection layers.","lead":"This paper simulates quantum time evolution of the lattice Schwinger model embedded into small error-detecting codes and shows that sparse rounds of error detection, combined with physics-based postselection, reduce errors in measured observables under near-term depolarizing noise. It also compares Iceberg and Hypercube code layouts and finds Iceberg codes more efficient when simulations are limited to a fixed number of shots.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's numerical gains are computed under depolarizing noise only, so the central claim depends on the untested assumption that coherent errors from non-FT rotation gadgets do not propagate into undetectable logical errors.","rationale":"The reader's weakest assumption identifies the same load-bearing condition: the paper's benefit relies on limited coherent error propagation from non-FT rotations, and the simulations use only depolarizing noise. I agree with this assessment. The paper is a careful circuit-level study: it explicitly constructs the encoded circuits, checks acceptance-rate scalings with independent larger-L data, and openly discusses saturation and resource trade-offs. These are real supporting elements. However, the central quantitative message compares systematic errors in observables, and those comparisons are performed entirely under a stochastic depolarizing noise model. The paper itself concedes that the non-FT time-evolution circuits can generate undetectable correlated logical errors, yet Appendix A only illustrates Pauli-error propagation through CNOT gates and never bounds the coefficient of such errors for the specific rotation gadgets. A coherent over-rotation or similar systematic error would not be represented by the depolarizing channel and could accumulate into the codespace, evading both stabilizer checks and postselection. Since the headline claim is about utility under realistic near-term noise, and realistic devices have coherent error components, this unquantified assumption is the most load-bearing concern. It does not invalidate the paper, but it does mean the central claim should remain conditional until the coherent-error sensitivity is tested. The reader already issued CONDITIONAL, so my recommendation is UNCHANGED.","tokens_in":43501,"tokens_out":5227,"duration_ms":53079,"concrete_test":"Run the L=2 [[4,2,2]]⊗2 simulations of Section III.B with the depolarizing channel augmented by a coherent over-rotation on every CNOT, e.g., apply exp(-i ε Z⊗Z/2) after each CNOT with ε chosen so the average gate infidelity equals p2=0.003 (and similarly for p2=0.001). Sweep nd=0,1,2,4 and compute the chiral condensate systematic error at t≈19, as in Fig. 12. If the error-detection benefit relative to the depolarizing-only baseline is reduced by more than the quoted uncertainties, the central claim is not robust to coherent non-FT error propagation. A complementary hardware check would be to run the same circuits on a device with characterized coherent error rates and compare acceptance-corrected systematic errors.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim is only as strong as the assumption, stated in Section II.A, that non-FT logical rotations used for time evolution have limited error propagation. The paper's own text says a non-FT rotation 'generates undetectable correlated two-qubit errors of the form of a logical operation,' and that non-FT gadgets can be used only 'provided the error propagation from non-FT gadgets is limited.' All numerical demonstrations are obtained with qiskit AerSimulator depolarizing noise (Section III.A), which is an incoherent, stochastic Pauli channel. Real near-term CNOT and rotation gates also have coherent systematic errors, e.g., over-rotations. A coherent error inside the non-FT gadgets in Figs. 8 and 9 can propagate into a weight-2 logical operator that commutes with the stabilizers, so it escapes both mid-circuit syndrome checks and final charge postselection, and it accumulates in the Trotter step count rather than being parametrically suppressed by detection. The paper never quantifies this propagation: Appendix A only illustrates Pauli errors through CNOT and provides no coefficient or bound for the specific rotation gadgets used in the simulations. If the undetectable logical-error contribution from coherent gate errors is comparable to the O(p2) floor claimed, the observed improvements in Figs. 12 and 20 could shrink below statistical significance. This is the load-bearing condition for the central claim because every simulated benefit is computed under a noise model that excludes it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses classical noisy simulation of small Schwinger-model instances to argue that sparse error detection in distance-2 codes can improve observable accuracy in near-term quantum simulations. The authors embed L=2 and L=4 staggered fermion systems into Iceberg codes [[N+2,N,2]] and compare with Hypercube codes [[2^N,N,2]], using depolarizing noise with p_2=0.001 and 0.003, fixed shot budgets, final charge-sector postselection, and a small number of mid-circuit stabilizer layers. They find that a modest number of error-detection layers reduces the systematic error in the chiral condensate and electric-field energy, that the improvement saturates as layers are added, that acceptance rates are predictable from gate counts and Hilbert-space dimensions, and that monolithic Iceberg encodings outperform Hypercube encodings in fixed-resource scenarios. They also provide gate-count scalings and extrapolated error-rate requirements for a 100-logical-qubit simulation.","tokens_in":43718,"tokens_out":9341,"duration_ms":85872,"significance":"Should the findings be robust, the paper provides a concrete, quantitative case for a 'sparse error detection' strategy: partial fault tolerance can be introduced with small overhead and without full FT rotation synthesis, improving observables in gauge-theory simulations. The main strengths are the level of detail (explicit circuits, noise parameters, shot budgets), the cross-size check in Fig. 14 where an L=2 fit predicts acceptance at L=4 and L=8, the injection-model cross-check of acceptance rates, and the honest enumeration of idealizations (no measurement errors, all-to-all connectivity, no post-processing mitigation). The paper does not claim a full error-correction threshold and explicitly identifies the non-FT rotation caveat, which helps the reader evaluate the scope. However, because the simulations are all depolarizing, the significance for real hardware remains conditional until coherent-error propagation in the non-FT gadgets is quantified.","major_comments":[{"comment":"The central claim is tested only under a depolarizing noise model, and the paper's own text concedes that non-FT rotations 'generate undetectable correlated two-qubit errors of the form of a logical operation' and that non-FT gadgets are usable only 'provided the error propagation from non-FT gadgets is limited.' Appendix A only illustrates Pauli-error propagation through a CNOT; it provides no coefficient or bound for the specific rotation gadgets in Figs. 8 and 9. Since coherent errors such as over-rotations can propagate into weight-2 logical operators that commute with the stabilizers and survive charge postselection, the gains shown in Figs. 12, 20, and 22 could be reduced or reversed on real hardware. The authors should add an explicit robustness test, e.g., injecting a coherent rotation error of the form theta -> theta(1+epsilon) into one non-FT gadget and comparing the accepted-ensemble bias with the depolarizing baseline, or else provide an analytic bound on the undetectable logical-error amplitude from those gadgets.","section":"Section II.A, Appendix A"},{"comment":"The paper claims an optimal density of error-detection layers, but the evidence is saturation of an empirical fit. Equation (15), f_chi = A exp(-gamma t/n_d) + (C-A), is asserted rather than derived and is stated to be invalid at late times; the fit parameters in Table I carry sizeable uncertainties (A=0.92(24), gamma=0.0133(46), C=0.545(7)), and the acceptance-rate cost is not folded into the same objective. The data support the statement 'improvement saturates', but they do not establish an operational optimum. Please either define an explicit cost function, such as the RMSE in Eq. (17), minimize it over n_d, or soften the optimality claim throughout the manuscript.","section":"Section IV.A, Eq. (15)"}],"minor_comments":[{"comment":"The noise model is described as isotropic depolarizing with p1 = p2/10 and p_rz = 0, but the exact depolarizing parameter used for single-qubit gates in qiskit's depolarizing_error is not stated; please specify the single-qubit depolarizing probability explicitly.","section":"Section III.A"},{"comment":"The injection-model curve in Fig. 27(a) is only shown for N=8; the caption should state this and quantify how well the single-N curve represents the other lattice sizes shown.","section":"Fig. 27"},{"comment":"The detection probabilities f1 ~ 0.90 - 0.15/k and f2 ~ 0.88 - 0.06/k are given without a stated range of validity in k and N; please specify the fitting range and the associated uncertainties.","section":"Section V.1, Eq. (27)"},{"comment":"The resource extrapolations quote p_shot,req and p_acc,req to two significant figures even though the underlying N>=20 points in Fig. 26 are labeled approximate; please present these as order-of-magnitude estimates or propagate the extrapolation uncertainty.","section":"Table II"},{"comment":"The manuscript would benefit from a code or data availability statement; the circuit diagrams are detailed, but exact reproduction of the qiskit simulations would be substantially easier with the scripts used to generate the noise model and gate counts.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is a careful simulation study, but the abstract's 'realistic noise rates' phrasing is broader than the depolarizing-only evidence. I would ask the authors to add a coherent-error propagation test or an analytic bound, and to soften the optimal-density claim if such a test is not possible. The companion hardware reference [88] may supply supporting evidence, but this manuscript should be self-contained enough for the central claim to be evaluated on its own."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one. It's a careful, honest simulation study, and the central result is believable within its stated scope. The paper compares Iceberg and Hypercube distance-two error-detecting codes on small Schwinger lattices and shows that a modest number of mid-circuit stabilizer measurements plus final charge postselection systematically reduce errors in the chiral condensate and gauge-field energy, at the cost of accepted shots. The new quantitative content is real: the saturation of improvement with detection-layer density, the partition trade-offs at L=4, the universal acceptance curve in x=p2 G_tot, and the L=2-based acceptance-rate prediction that is confirmed at L=4 and L=8. That last point is a genuine strength—it's a parameter-free prediction test, not a fit to the target observable.\n\nThe simulations are carefully specified: explicit circuits, noise parameters, gate counts, and fit forms. The injection model in Section V.2 is a nice independent check on the acceptance rates. The paper is also honest about saturation, floors, and extrapolation uncertainties. Citation pattern is normal for this group; self-citations point to directly relevant prior work, not padding.\n\nThe main soft spot is the noise model. Everything runs under depolarizing noise, no measurement errors, all-to-all connectivity. The time-evolution rotations are explicitly non-FT, and the paper acknowledges they can generate undetectable correlated logical errors. Section II.A says the method works 'provided the error propagation from non-FT gadgets is limited,' but Appendix A only illustrates Pauli propagation through CNOT; it never bounds the coefficient for the actual rotation gadgets in the simulations. So the abstract's 'realistic noise rates' is really 'stochastic Pauli noise at realistic rates.' That is a meaningful gap. A coherent over-rotation inside those gadgets could produce exactly the kind of undetectable weight-2 logical error that postselection can't catch, and the benefit could shrink. The stress-test note is right to flag it. It's not fatal, because the paper doesn't claim to model coherent errors, but it should be scoped clearly or addressed with coherent-noise runs before people use the numbers as a resource guide. Also, no code or data is shipped, so nothing is independently reproducible, and the 100-qubit projections are fit-based extrapolations.\n\nWho is this for: groups planning near-term lattice-gauge-theory simulations on all-to-all connected hardware, and people comparing error-detecting code families. It deserves a serious referee. My recommendation: send it out, but ask for coherent-noise simulations or an analytic bound on undetectable logical error, and for code/data release.","headline":"A careful, honest simulation study showing sparse error detection helps for Schwinger-model observables under depolarizing noise; the main caveat is that coherent errors from non-FT rotations are never quantified.","tokens_in":44380,"tokens_out":3173,"would_cite":true,"duration_ms":30264,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Under realistic near-term noise rates, adding a small number of error-detection layers to distance-two encoded quantum simulations of the Schwinger model reduces the systematic error of observables, with an optimal density beyond which no…","keywords":["quantum error detection","Iceberg codes","lattice gauge theory","Schwinger model","sparse stabilizer checks","postselection","quantum simulation","fault tolerance"],"falsifier":"Run the [[10,8,2]]-encoded L=4 Schwinger evolution on a device with all-to-all connectivity at a two-qubit error rate near p2=0.001, with 4 and 8 stabilizer layers and at least $10^{6}$ shots; the claim predicts the systematic error in the chiral condensate drops steadily with layer count and saturates near eight layers, so observing no improvement over unencoded charge-postselected results, or a rise in error as layers are added, would refute it.","tokens_in":43213,"feed_emoji":"⚛️","tokens_out":9908,"duration_ms":84914,"temperature":0.7,"pith_summary":"This paper argues that a small amount of quantum error detection can make near-term quantum simulations of lattice gauge theories more accurate, even when the circuits are not fully fault-tolerant. Using distance-two codes from the Iceberg and Hypercube families to encode the Schwinger model, the authors find that adding a few mid-circuit stabilizer measurements and postselecting on charge conservation removes the leading-order errors in observable estimates. The improvement appears at realistic two-qubit noise rates and comes at the price of discarding an exponentially shrinking fraction of runs. If the claim holds, minimal fault-tolerant components could extend the reach of quantum simulations before full error correction is available.","feed_headline":"Sparse error checks lift accuracy in noisy gauge-theory simulations","feed_subtitle":"A few mid-circuit stabilizer measurements plus charge postselection cut systematic error while discarding some shots.","key_machinery":"The load-bearing object is the distance-two stabilizer code family [[N+2,N,2]], called Iceberg codes, in which N logical qubits are encoded into N+2 physical qubits and every codeword is a superposition of two complementary bit strings, so the stabilizers S_Z=$Z^{{⊗(N+2)}}$ and S_X=$X^{{⊗(N+2)}}$ detect any single-qubit error. The protocol intersperses fault-tolerant measurements of these stabilizers at sparse intervals during a Trotterized time evolution whose rotation gates are deliberately not fault-tolerant, and then postselects on both the stabilizer outcomes and conservation of electric charge in the final measurement. The encoding does the work: it enlarges the Hilbert space so that pairs of O(p) errors that would conspire to re-enter the codespace are pushed to an O($p^{2}$) floor, while the charge postselection removes O(p) charge-violating errors for free. The [[2^N,N,2]] Hypercube family is used as a comparison, with its transversal gates but higher shot-rejection rates.","core_discovery":"The central discovery is that sparse, non-fault-tolerant error detection has genuine utility for quantum simulation of gauge theories in the near term. Embedding the lattice Schwinger model into [[N+2,N,2]] Iceberg code blocks, where each logical qubit is a GHZ-type superposition spread across physical qubits, and inserting a small number of fault-tolerant stabilizer measurements during Trotter evolution, followed by postselection onto the charge-zero sector, systematically reduces the systematic error of local observables such as the chiral condensate and electric-field energy. The reduction is not monotone in the number of detection layers: for a fixed evolution time and error rate there is an optimal density of layers, and beyond it the error saturates at a code-dependent floor while the accepted ensemble continues to shrink. For the L=2 system, one or two layers at p2=0.003 already improve on unencoded charge-postselected results; for L=4, the recovered fraction of the systematic error follows fχ(t,nd)=A $e^{{-γt/n_d}}$+(C-A) and saturates near 0.55 for eight layers. The acceptance rate collapses onto a universal curve in the resource variable x=p2G_tot, with a floor-subtracted crossing at x*=6.23±0.16, which underpins extrapolations to larger systems.","pith_inferences":["The universal acceptance curve in the resource variable x=p2G_tot suggests a practical rule of thumb: for monolithic Iceberg codes, plan around x*≈6.2 as the useful noise budget per evolution, independent of system size for N≥6.","The optimal-layer saturation implies that pre-production tuning can fix the detection-layer density once for a given Hamiltonian, error rate, and target time, rather than requiring a per-observable optimization.","Combining sparse error detection with standard error-mitigation techniques, which the paper explicitly leaves for future work, could push the useful time horizon further because detection removes leading-order errors that mitigation would otherwise have to extrapolate away.","The same sparse-detection protocol should be testable in other charge-conserving lattice theories, where Gauss's law supplies a symmetry-based postselection that costs no extra gates."],"forward_implications":["For the small Schwinger-model systems studied, a single mid-circuit stabilizer layer plus final charge postselection reduces systematic error in the chiral condensate and electric-field energy compared with unencoded postselected evolution at p2=0.003.","There is an optimal density of error-detection layers for a given evolution time and noise rate; beyond that density, accuracy saturates while the accepted ensemble continues to shrink.","In fixed-shot-resource comparisons at L=4, the monolithic [[10,8,2]] block is best at low shot counts, while the [[6,4,2]]⊗[[6,4,2]] partition is favored at high shot counts.","Hypercube encodings match Iceberg accuracy but reject more shots for the same evolution, so they underperform in resource-constrained scenarios.","Extrapolations put the useful operating point for a 100-logical-qubit monolithic Iceberg simulation at two-qubit error rates near 10^-5 to 10^-6 depending on the number of Trotter steps."],"supporting_citations":[{"why":"Defines the distance-two Iceberg code family [[N+2,N,2]] used for all encoded simulations.","marker":"[59–64]"},{"why":"Supplies the axial-gauge Schwinger-model Hamiltonian and the large-scale simulation context that set the resource targets.","marker":"[67, 68]"},{"why":"Provides the Gauss's-law check circuits that make charge postselection possible.","marker":"[39, 40]"},{"why":"Demonstrates that encoded circuits with non-FT rotations plus FT gadgets can beat unencoded ones, the approach adopted here.","marker":"[8, 88]"},{"why":"The classical noisy-simulation tool used to generate all numerical results.","marker":"[103]"},{"why":"Defines the [[8,3,2]] Hypercube code and its transversal gate structure used in the comparison.","marker":"[114]"}],"fun_headline_variants":["Sparse error checks push gauge-theory sims past noise floor","Iceberg codes make noisy gauge simulations more accurate","Sparse error detection cuts systematic error in quantum sims","Fewer shots, better accuracy: sparse stabilizer checks in gauge theories","Postselection plus sparse checks beat noise in Schwinger sims"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole benefit rests on the assumption that the non-fault-tolerant rotation gates used for time evolution spread errors only mildly, so that the undetectable errors they create do not outweigh what the stabilizer checks remove; this is tested only under depolarizing noise with no measurement errors and all-to-all connectivity.","fun_headline_variants_meta":{"raw":{"variants":["Sparse error checks push gauge-theory sims past noise floor","Iceberg codes make noisy gauge simulations more accurate","Sparse error detection cuts systematic error in quantum sims","Fewer shots, better accuracy: sparse stabilizer checks in gauge theories","Postselection plus sparse checks beat noise in Schwinger sims"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001491,"raw_usage":{"total_tokens":6028,"prompt_tokens":1031,"completion_tokens":4997,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":647,"completion_tokens_details":{"reasoning_tokens":4912}},"tokens_in":647,"tokens_out":4997,"duration_ms":40068,"temperature":1.0,"reasoning_tokens":4912,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:54:55.909482+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the [[10,8,2]]-encoded L=4 Schwinger evolution on a device with all-to-all connectivity at a two-qubit error rate near p2=0.001, with 4 and 8 stabilizer layers and at least $10^{6}$ shots; the claim predicts the systematic error in the chiral condensate drops steadily with layer count and saturates near eight layers, so observing no improvement over unencoded charge-postselected results, or a rise in error as layers are added, would refute it.","supporting_citations":[],"review_version":1}