{"id":"4ed78ea8-9654-4e24-9dfe-bdf81e9d8ec2","arxiv_id":"2506.20125","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"On IBM quantum processors, self-mitigation corrects noisy Trotterized quench dynamics of Heisenberg spin chains (up to 104 qubits, over 3,000 CNOT gates) more accurately and stably than zero-noise extrapolation.","lead":"Researchers ran noisy quantum simulations of Heisenberg spin chains with up to 104 qubits on IBM machines, and found that a correction method called self-mitigation stays accurate as circuits grow, while the standard extrapolation approach degrades. The work is a step toward using today's noisy quantum computers for many-body physics before fault-tolerant hardware arrives.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Large-scale ZNE comparison reuses the N=20 extrapolation curve, so the claim that SM outperforms ZNE at 84-104 qubits is not established by the reported data.","rationale":"The reader's weakest assumption (test/target noise-factor equality) is a genuine unverified premise, but it is partially supported by the N=20 validation and by the visual agreement of SM with MPS at large scale. The ZNE comparison flaw is not speculative: the paper openly reuses a fit from a different system size, which is not a valid ZNE implementation. Because the headline claim emphasizes SM's advantage over ZNE at utility scale, this is the most load-bearing threat. A concrete refit test is straightforward and can settle it. The verdict remains conditional, as the paper still contributes a useful mitigation scheme and small-scale validation, but the large-scale comparative claim requires revision.","tokens_in":20144,"tokens_out":11545,"duration_ms":123967,"concrete_test":"Rerun the ZNE analysis for N=104 OBC and N=84 PBC using only the noise-amplified expectation values measured at those system sizes (folding factors 1, 3, and 5), fit the same functional form (e.g., polynomial) used at N=20, and recompute the mean absolute error relative to the MPS-TDVP benchmark. If the refit ZNE error remains above SM (0.02556 OBC, 0.02832 PBC), the conclusion survives this check; if it falls below or near SM, the paper's claim that SM outperforms ZNE at utility scale is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2 states: 'For the ZNE, we apply the same extrapolation fitting curve used in the N=20 cases.' Standard ZNE fits a noise-amplification model to expectation values measured at the same system size; reusing the N=20 fit at N=104/84 pre-assumes the noise-vs-folding relationship is size-independent, which the paper itself argues breaks down. The large-scale comparison therefore pits SM against a deliberately mismatched ZNE, and the reported 43.9-48.0% error reduction at scale is an artifact of this choice. Since the central claim includes 'clearly outperforming both ZNE and the baseline' across all tested sizes, this design flaw directly undermines the headline conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a self-mitigation (SM) error mitigation method for Trotterized time evolution, extending the method of Ref. [12] to optimized second-order Trotterization, and applies it to quench dynamics of the Heisenberg XXZ spin chain on IBM quantum hardware. On N=20 systems, SM applied on top of TREX+DD+PT is compared with ZNE and the baseline, showing lower mean absolute errors in staggered magnetization. The method is then deployed at N=104 (OBC) and N=84 (PBC) with 3,246 and 2,646 CNOT gates, respectively, where the paper reports that SM maintains errors near the N=20 values while ZNE errors increase. The paper also demonstrates a randomized-measurement protocol for Rényi entropy on an IBM Heron processor.","tokens_in":20139,"tokens_out":7080,"duration_ms":63817,"significance":"If the central claim held, the paper would provide a valuable, resource-efficient QEM technique that avoids ZNE's extrapolation overhead and remains stable at utility-scale qubit counts. Strengths include the clean N=20 benchmark against exact simulation, the transparent reporting of 10 repetitions and 100,000 shots per circuit, and the validation of the MPS-TDVP reference at N=20. However, the large-scale comparison is compromised by an unfair ZNE baseline, and the SM noise-factor-transfer assumption is only validated at small size, so the headline claim of 'clearly outperforming both ZNE and the baseline' across all tested sizes is not yet established.","major_comments":[{"comment":"The claim that SM outperforms ZNE at N=104/84 is not supported by the reported data because the ZNE comparison is deliberately mismatched. The manuscript states 'For the ZNE, we apply the same extrapolation fitting curve used in the N=20 cases,' yet §3.4 argues that ZNE's assumption of a consistent noise model 'often breaks down when the number of qubits changes.' Reusing the N=20 curve at N=84/104 therefore tests ZNE under a known-failing condition, and the reported 43.9–48.0% error reductions are an artifact of this handicap. A fair comparison requires fitting the ZNE curve to measurements at each system size, or at least reporting both a re-fit ZNE and the transferred-curve ZNE. Without this, the headline conclusion that SM 'clearly outperforming both ZNE and the baseline' across all tested sizes is not established.","section":"§4.2, Figs. 11–13, Table A3"},{"comment":"The self-mitigation correction factor 1/(1−p) rests on the assumption that the noise factor p measured on the test circuit equals that of the target circuit. The test circuit with the (dt, dt, −dt, −dt) sequence still has a vanishing center layer (the U_j(0) layer), so it is not gate-identical to the target; the paper acknowledges this only by assertion ('we assume the discrepancy ... is negligible'). The equality is checked empirically only at N=20. Because a size-dependent bias in p would propagate multiplicatively into every reported SM value, the N=20 check is not sufficient to justify the large-scale results. Please provide a quantitative test of the size dependence of p (e.g., compare p extracted at N=20 and N=104, or show that the reported errors are insensitive to a ±20% perturbation of p).","section":"§3.5, Figs. 2–3"}],"minor_comments":[{"comment":"The names 'ibm yonsei' and 'ibm marrakesh' should be formatted consistently (e.g., IBM Yonsei, IBM Marrakesh) and the processor generation should be stated uniformly for each device used.","section":"§4, hardware identification"},{"comment":"The caption 'Averaged staggered magnetization of ten repetition' should read 'ten repetitions'; similar grammar issues appear in a few other places.","section":"Table A1 caption"},{"comment":"The axis labels in Figure 16 appear corrupted by unicode escape sequences (e.g., '/uni00000013'); the figure should be regenerated with correct font encoding so that 'Rényi Entropy, S(2)' is readable.","section":"Figure 16"},{"comment":"The typo 'entanglment' should be corrected, and the number of shots per random-unitary circuit in the RM protocol should be stated explicitly.","section":"§5.2"},{"comment":"The product notation in Eq. (3) is ambiguous: the second product is written as ∏_{n=M}^1, which is nonstandard; please clarify the ordering of factors.","section":"Eq. (3)"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be of interest to the quantum error mitigation community. The main risk is the ZNE comparison in §4.2; the authors should be asked to re-fit ZNE at each system size or clearly label the transferred-curve comparison as a stress test rather than a head-to-head benchmark. The 'quantum utility' phrasing in the title and abstract is stronger than the evidence supports, since the target problem remains classically simulable with MPS-TDVP."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a useful, honestly limited engineering study of self-mitigation (SM) for Trotterized Heisenberg quench dynamics. The N=20 validation against exact simulation is clean, and scaling to 84/104 qubits is worth reporting. But the headline claim that SM outperforms ZNE at scale is not established, because the large-scale ZNE comparison reuses the N=20 extrapolation curve.\n\nWhat is genuinely new: adapting the SM test circuit to optimized second-order Trotterization with a (dt,dt,-dt,-dt) sequence, the 20-to-104 qubit scaling study with more than 3,000 CNOTs, and the randomized-measurement Rényi entropy results combined with QMP. The paper reports 10 repetitions, 100,000 shots, and full tables, which is more reproducible than most hardware papers. The authors also explicitly flag the central assumption in Section 3.5, which deserves credit. The citation pattern looks proper: SM is traced to Ref. [12] and the depolarizing framework to Ref. [46].\n\nSoft spots, in order of severity. First, Section 4.2 states: 'For the ZNE, we apply the same extrapolation fitting curve used in the N=20 cases.' Standard ZNE should fit at each system size; reusing the small-system fit handicaps ZNE and makes the 43.9-48.0% error reduction at scale partly an artifact. This is the load-bearing flaw for the comparison. Second, the SM noise factor p is measured on a test circuit that is not gate-identical to the target: the middle U(0) layer vanishes, and half the angles are reversed. The authors assume the discrepancy is negligible, but that is checked only at N=20. If p drifts at 84/104 qubits, the reported SM values are systematically distorted. Third, the MPS-TDVP benchmark at bond dimension 1000 has no convergence check; it is probably fine for these short times, but 'probably' is not a benchmark. Fourth, the entropy measurement reports 1.73% agreement without error bars and omits the subsystem size, with only 60 random unitaries. Finally, the 'quantum utility' framing overstates the case: staggered magnetization on a 1D chain is classically easy, and the paper's own MPS benchmark confirms it.\n\nRecommendation: send it to peer review. A good referee can push for a re-fit ZNE baseline, a p-transfer check, and MPS convergence checks. I would cite the N=20 comparison and the test-circuit construction if I were writing about QEM for Trotterized evolution.","headline":"Useful SM validation at N=20 and a plausible scaling study, but the large-scale ZNE comparison reuses the N=20 fit and the p-transfer assumption is only checked at small size.","tokens_in":20861,"tokens_out":3209,"would_cite":true,"duration_ms":31641,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single noise correction from a forward-backward test circuit keeps 104-qubit spin-chain simulations accurate and beats zero-noise extrapolation.","keywords":["quantum error mitigation","self-mitigation","zero-noise extrapolation","Trotterization","Heisenberg XXZ spin chain","quantum quench dynamics","staggered magnetization","entanglement entropy"],"falsifier":"At a system size where the ideal output of the target circuit is known at early times, for example the 20-qubit case with exact simulation, extract the depolarizing noise factor of the test circuit and of the target circuit separately, either by randomized benchmarking on the same gate layers or by comparing corrected early-time values, and check whether they agree within the reported error bars; a disagreement larger than the error bars would mean the $1/(1-p)$ correction is systematically biasing the staggered-magnetization results.","tokens_in":19775,"feed_emoji":"⚛️","tokens_out":11417,"duration_ms":102640,"temperature":0.7,"pith_summary":"The paper tries to establish that a normalization-based error-mitigation scheme, called self-mitigation, keeps time evolution on today's noisy superconducting hardware accurate even when the simulation is Trotterized into thousands of two-qubit gates and run on 104 qubits. The scheme runs a companion test circuit that performs half of the Trotter steps forward and half backward, so its ideal output is the known initial state; the decay of the expectation value in that test circuit gives a depolarizing noise factor $p$, and the target circuit is corrected through $\\langle O\\rangle = \\langle O\\rangle/(1-p)$. For quench dynamics of the Heisenberg XXZ spin chain, the paper reports mean absolute errors in staggered magnetization around 0.026 at 104 qubits with self-mitigation versus around 0.049 with zero-noise extrapolation, and errors that stay at 20-qubit levels rather than growing. A sympathetic reader would care because it points to an extrapolation-free way to get utility-scale accuracy for time-dependent many-body problems before fault-tolerant quantum computers arrive.","feed_headline":"Self-mitigation holds accuracy at 104 qubits in spin-chain quench","feed_subtitle":"A forward-backward test circuit estimates device noise in place of curve fitting, keeping errors flat as qubits grow","key_machinery":"The load-bearing object is the self-mitigation test circuit: the same optimized second-order Trotter circuit as the target, but with the first half of the steps executed with $+dt$ and the second half with $-dt$. Because forward evolution followed by the exact reverse evolution returns the initial state, the ideal expectation value of any observable in the test circuit is known in advance; the ratio between the measured noisy value and the known ideal value yields the depolarizing noise factor $p$, and the target circuit's noisy expectation value is divided by $1-p$. The paper chooses the time sequence $(dt, dt, \\dots, -dt, -dt)$ rather than alternating signs so that only the middle layer of $U_j(\\vec 0)$ identity gates drops out, keeping the test circuit structurally near the target. The work this machinery does is to replace extrapolation in noise strength with a single multiplicative correction; the whole argument rests on the test circuit and target circuit suffering the same noise factor.","core_discovery":"On its own terms, the discovery is that a depolarizing-noise correction factor measured from a structurally similar test circuit stays valid as the system grows. For the XXZ model with open and periodic boundary conditions, the authors stack readout twirling, dynamical decoupling, and Pauli twirling, then add self-mitigation: a test circuit with the time sequence $(dt, dt, \\dots, -dt, -dt)$ whose ideal final state is the initial Néel state. The measured expectation value in that test circuit gives $p$, and the target circuit is corrected by $1/(1-p)$. Across $N=20$, $N=84$ (periodic), and $N=104$ (open) qubits, the corrected staggered-magnetization curves track the classical tensor-network benchmark closely; mean absolute errors are 0.02760 (20, open), 0.02832 (84, periodic), and 0.02556 (104, open), versus 0.04106, 0.05044, and 0.04917 for zero-noise extrapolation. The claim is that self-mitigation's circuit-aware normalization captures the dominant noise without extrapolation, while ZNE's fitted curve, derived at 20 qubits, degrades at larger scale.","pith_inferences":["Inference: the same forward-backward construction should transfer to other time-symmetric protocols, such as Loschmidt echoes, echo-type diagnostics, or out-of-time-order correlator measurements, where the ideal output is known; the transfer would still require the test/target noise-match assumption to hold.","Inference: because self-mitigation applies one global depolarizing factor, spatially non-uniform or correlated noise should reveal itself as a late-time systematic drift in the corrected curves; comparing self-mitigated results against independent benchmarks across circuit depths would quantify this sensitivity.","Inference: the method's applicability is limited to cases where the initial-state expectation value is known, so a practical protocol would use self-mitigation for Trotterized quench observables and fall back to zero-noise extrapolation for general circuits; the paper itself notes ZNE's broader applicability but not this hybrid."],"forward_implications":["Correcting by a single measured noise factor instead of extrapolating from amplified circuits keeps the mean absolute error in staggered magnetization near 0.026 at 104 qubits, essentially unchanged from the 20-qubit value.","Zero-noise extrapolation's error grows by about 20 percent from 20 to 104 qubits when its fitting curve is fixed at the small-system size, while self-mitigation does not require such a curve.","Self-mitigation adds exactly one test circuit per observable, whereas the zero-noise extrapolation schedule used here adds two noise-amplified circuits, so the scheme is more resource-efficient in both circuit count and shot overhead.","Combining the same mitigation stack with randomized measurement protocols and circuit parallelization yields Rényi entanglement-entropy estimates within a few percent of classical simulation, extending the mitigation benefit beyond simple observables."],"supporting_citations":[{"why":"Introduces self-mitigation, the protocol that this paper extends to optimized second-order Trotterization.","marker":"[12]"},{"why":"Supplies the depolarizing-noise mitigation relation between noisy and corrected expectation values that self-mitigation specializes.","marker":"[46]"},{"why":"Provides the optimized second-order Trotter circuit used as the target evolution for the XXZ chain.","marker":"[14]"},{"why":"One of the two founding references for zero-noise extrapolation, the main method self-mitigation is compared against.","marker":"[5]"},{"why":"The companion founding reference for zero-noise extrapolation as a noise-scaling strategy.","marker":"[6]"},{"why":"Supplies the local unitary gate-folding schedule with scaling factors 1, 3, and 5 used to implement zero-noise extrapolation.","marker":"[13]"},{"why":"Defines the twirled readout error extinction method included in the baseline mitigation stack.","marker":"[37]"},{"why":"Establishes Pauli twirling, which converts coherent errors into stochastic Pauli errors and is part of the baseline stack.","marker":"[42]"},{"why":"Provides the time-dependent variational principle classical simulation used as the benchmark for the 84- and 104-qubit results.","marker":"[50]"},{"why":"Provides the direct exact state-vector simulation used to validate the 20-qubit reference and the entropy results.","marker":"[47]"}],"fun_headline_variants":["New self-mitigation tames noise at 104 qubits","Self-mitigation beats extrapolation on 104 qubits","Quench dynamics preserved by circuit-aware noise fix","No extrapolation needed: self-mitigation at 104 qubits","Self-mitigation outshines ZNE at 104 qubits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the noise measured in the forward-backward test circuit is the same as the noise in the actual forward simulation, even though the test circuit leaves out a whole layer of identity gates; the paper states this as an assumption and checks it only at 20 qubits.","fun_headline_variants_meta":{"raw":{"variants":["New self-mitigation tames noise at 104 qubits","Self-mitigation beats extrapolation on 104 qubits","Quench dynamics preserved by circuit-aware noise fix","No extrapolation needed: self-mitigation at 104 qubits","Self-mitigation outshines ZNE at 104 qubits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1538,"prompt_tokens":991,"completion_tokens":547,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":459}},"tokens_in":607,"tokens_out":547,"duration_ms":5649,"temperature":1.0,"reasoning_tokens":459,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:22:13.151106+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"At a system size where the ideal output of the target circuit is known at early times, for example the 20-qubit case with exact simulation, extract the depolarizing noise factor of the test circuit and of the target circuit separately, either by randomized benchmarking on the same gate layers or by comparing corrected early-time values, and check whether they agree within the reported error bars; a disagreement larger than the error bars would mean the $1/(1-p)$ correction is systematically biasing the staggered-magnetization results.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the depolarizing-noise mitigation relation between noisy and corrected expectation values that self-mitigation specializes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the optimized second-order Trotter circuit used as the target evolution for the XXZ chain."}],"review_version":2}