{"id":"a391121f-fb7a-40a2-b750-f42f02e7f40f","arxiv_id":"2505.10820","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"MLXEB estimates circuit fidelity for large quantum devices using particle-number-conserving random circuits whose ideal output distribution can be classically simulated in a reduced Hilbert space.","lead":"The paper introduces a new way to benchmark quantum computers with more than about 50 qubits by using circuits that conserve particle number, which makes the ideal output easy to compute classically. The method, called MLXEB, estimates overall circuit fidelity from measured bitstrings, and numerical tests on lattices up to 196 qubits suggest it works when noise is low.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Scalability beyond 100 qubits rests on unverified fixed-angle Porter-Thomas convergence; Sections IV B and V leave this premise open.","rationale":"The paper's central claim is that FMLXEB,n estimates circuit fidelity for >100-qubit devices using particle-number-conserving circuits. I checked the estimator under a simple global depolarizing model: if the ideal distribution is Porter-Thomas on the n-particle subspace and the noisy state is (1-p)|ψ⟩⟨ψ| + p I/D, then FMLXEB,n → 1-p, matching the Uhlmann fidelity F in Eq. (4), provided D_n p_n follows the rescaled exponential. The estimator's prefactor ns/Ns plays the correct compensating role, so the mathematics of Eq. (11) is internally sound. The fragile point is therefore the distributional premise, and it is most fragile for the fixed-angle ensemble that the paper actually recommends for practical noisy benchmarking. The numerical evidence for Porter-Thomas convergence in the fixed-angle case stops at N=100 (Section IV B), while Section V concedes that neither t-design membership nor the scaling relation (13) holds consistently for fixed-angle circuits. The only noisy validation is a single N=64, n=1 setup without explicit error bars (Fig. 8). This does not disprove the method; it identifies a concrete missing validation in the advertised regime. The reader's conditional verdict already captures this, so no verdict change is needed. The proposed test would settle whether the >100-qubit extrapolation is justified.","tokens_in":13939,"tokens_out":13583,"duration_ms":148233,"concrete_test":"Run noiseless fixed-angle (α1=π/8, α2=0) MLXEB circuits for N=196 and N=256, with n=1,2,3 and the center-initialized particle, at depths Nd≈20, 40, 80, and 160. For at least 50 independently sampled circuits, histogram q=Dn p_n(x) over all outcomes and compute (i) the second moment Dn⟨p_n²⟩, which equals ≈2 for Porter-Thomas, and (ii) a Kolmogorov-Smirnov or chi-square comparison against PT(q)=e^{-q}, with statistical error bars. If the fixed-angle data at N=196/256 depart from Porter-Thomas beyond sampling error at the depths used in Fig. 8, the normalization behind Eq. (11) fails in the advertised >100-qubit regime. Additionally, attempt to collapse the fixed-angle FMLXEB data using Eq. (13) with a=0.3; absence of collapse would confirm that the Section V admission also blocks depth prediction for the practical fixed-angle configuration.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The practical protocol for noisy devices uses fixed-angle XXZ gates (α1=π/8, α2=0) with a center-initialized particle because Fig. 5 shows faster convergence. But the central estimator FMLXEB,n in Eq. (11) is normalized only if the ideal output distribution is the rescaled Porter-Thomas distribution on the n-particle subspace. For the fixed-angle ensemble, Porter-Thomas convergence is numerically checked only for N≤100 (Section IV B), and Section V explicitly leaves open both the t-design membership and a scaling law for this ensemble, stating that the scaling relation in Eq. (13) does not appear to hold consistently across the range of N considered. Since the advertised target is >100 qubits and the only noisy validation is a single N=64, n=1, α1=π/8 setup (Fig. 8), the load-bearing premise that fixed-angle circuits reach Porter-Thomas at the depths and scales used by MLXEB is unsupported. If the fixed-angle ensemble deviates from Porter-Thomas above N=100, or reaches it only at depths where depolarizing noise suppresses the signal, FMLXEB no longer tracks the fidelity F defined in Eq. (4). This is not an internal inconsistency, but an extrapolation beyond the supplied numerical evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes MLXEB, a modified linear cross-entropy benchmark for particle-number-conserving (U(1)-symmetric) random quantum circuits. By restricting the Hilbert space to a fixed particle-number sector of dimension D_n = C(N,n), the authors argue that the ideal output distribution can be simulated classically for N>100 when n=O(1). The estimator F_MLXEB,n in Eq. (11) is constructed so that, when the ideal distribution in the n-particle sector is the rescaled Porter-Thomas distribution, it should equal the circuit fidelity F. Numerical tests are reported for noiseless circuits up to N=196, for two gate-angle settings (random and fixed α1=π/8, α2=0), showing convergence of F_MLXEB to 1 with depth and a data collapse based on Eq. (13). Under depolarizing noise, a single N=64, n=1 fixed-angle configuration is reported to match the true fidelity for p_noise ≤ 2×10^-3. The paper is candid about open questions, including the t-design class of the fixed-angle ensemble and the absence of a consistent scaling relation for fixed angles.","tokens_in":14119,"tokens_out":14538,"duration_ms":137621,"significance":"If the claims hold, MLXEB would be a useful extension of LXEB to non-Clifford circuits on devices beyond 50 qubits, a gap the paper identifies correctly. The strengths include: the fidelity estimator is benchmarked against an independently computed true fidelity rather than against itself; the noiseless random-angle results show a scaling collapse over system sizes and particle numbers; and the use of particle-number conservation to enable classical simulation is sound. The paper also explicitly flags the main caveats, which is helpful. However, the practical fixed-angle variant—the one recommended for shallow noisy circuits—is validated for Porter-Thomas behavior only up to N=100 and for fidelity tracking only at N=64, so the headline scalability claim is only partially supported.","major_comments":[{"comment":"The fixed-angle circuit ensemble (α1=π/8, α2=0) is the variant used in the noisy benchmark and recommended for hardware, but its convergence to the rescaled Porter-Thomas distribution is verified only for N=36, 64, and 100 at depth Nd=160 (Fig. 2), and Section V explicitly states that the scaling relation in Eq. (13) does not hold consistently for this ensemble and that its t-design class is open. Since Eq. (11) is normalized by D_n and is valid only if the ideal n-particle distribution is the rescaled Porter-Thomas distribution, the extrapolation of the fixed-angle protocol to devices with more than 100 qubits—which appears in the abstract and conclusion—is unsupported. I request either a numerical test of fixed-angle Porter-Thomas convergence for N>100 (at least for n=1) or a revision that restricts the practical scalability claim to the random-angle protocol and states the fixed-angle protocol's validity as demonstrated only up to N=100.","section":"§IV B, §V"},{"comment":"The central noisy validation claims that F_MLXEB,1 agrees with the true circuit fidelity for 1.0×10^-3 ≤ p_noise ≤ 2.0×10^-3, and that the deviation at p_noise ≥ 3.0×10^-3 is due to sampling error scaling as O(1/sqrt(N_s N_c)). No error bars or confidence intervals are described, and the text does not quantify the variance of F_MLXEB under the Porter-Thomas model. Given that the asymptotic standard deviation of an LXEB-type estimator can be estimated as sqrt(2/n_s) (roughly 1.4×10^-3 for the total samples used here), the claimed sampling-error explanation for discrepancies at p_noise = 3×10^-3 needs a quantitative check rather than a qualitative statement. Please add error bars (for example, bootstrap over the N_c=10 circuit instances) and, if the discrepancy persists, discuss possible bias from the n_max=3 truncation.","section":"§IV D, Fig. 8"},{"comment":"The finite-size scaling relation in Eq. (13) is introduced with an exponent 1/(n+a) whose form is motivated heuristically, with a=0.3 chosen to achieve a collapse, and the stretched-exponential parameters in Eq. (15) are fit to the same data that are then shown as collapsed. The relation is used in the conclusion to predict the circuit depth required for devices larger than those simulated, but no held-out predictive test is reported. If the scaling relation is meant to guide experimental design for N>196, please validate it by fitting on a subset of (N,n) pairs and predicting held-out data, or clearly label it as an empirical observation for the studied sizes.","section":"§IV C, Eq. (13)"},{"comment":"The truncation strategy n_max=3, used for the principal N=64, n=1 noisy simulation, is validated in Fig. 6 only for N=16, n=3 by comparison with the full Hilbert space. The theoretical argument that leakage to sectors with |δn|>2 is O(p_noise^3) is reasonable, but since this truncation underlies the only fidelity-tracking demonstration, it would strengthen the paper to verify at a larger system size where full simulation is still possible (e.g., N=20 or 24 with n=1) that the observable F_MLXEB is insensitive to the truncation cutoff, or to provide an explicit bound on the induced error.","section":"§IV D, Fig. 6"}],"minor_comments":[{"comment":"The word 'consistenty' in the first paragraph of Section V should be 'consistently'.","section":"§V"},{"comment":"The text says Porter-Thomas convergence is confirmed for N=36 to 100, but the abstract and conclusion later refer to validation up to N=196; it should be stated explicitly that the N=196 validation is for noiseless F_MLXEB convergence (Fig. 3) and not for a direct Porter-Thomas histogram comparison.","section":"§IV B, §IV C"},{"comment":"The sentence 'validating the analytical expression used in Eq. (6)' appears to be a typo for Eq. (16), since Eq. (6) defines the particle-number sector G_n rather than the predicted-fidelity formula.","section":"§IV D"},{"comment":"The noise model applies a two-qubit Pauli selected uniformly from the 15 non-identity Paulis with probability p_noise; this is a valid depolarizing channel but should be stated explicitly as the convention used, to avoid confusion with the more common convention where each non-identity Pauli occurs with probability p_noise/15.","section":"§IV A"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The paper is honest and technically sound in what it directly demonstrates, and the main issue is an extrapolation gap rather than an internal inconsistency. I would not reject on the fixed-angle scaling issue alone, but the authors should be asked to either supply the missing fixed-angle Porter-Thomas data at N>100 or explicitly narrow the claims. The statistical-error discussion in §IV D also needs to be substantiated before publication. The novelty relative to Refs. [38,39] is moderate, but the MLXEB estimator and its noisy validation are new enough for a methods-oriented journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The new thing here is the MLXEB estimator in Eq. (11), with the ns/Ns prefactor that compensates for post-selecting on the n-particle sector. The idea is simple and natural: restrict random circuits to conserve particle number with n=O(1), the effective Hilbert space shrinks to polynomial size, classical simulation of circuits with 100+ qubits becomes feasible, and LXEB-style fidelity estimation still works. I am not aware of this exact combination in the prior LXEB literature, and the numerical work in Section IV is honest and fairly thorough. The noiseless convergence to F=1, the Porter-Thomas histograms for N=36 to 100, and the low-noise agreement with true fidelity in Fig. 8 all support the central claim. The authors also deserve credit for stating clearly what they did not establish: Section V admits that the fixed-angle ensemble's t-design membership is open and that the scaling relation does not hold consistently for fixed angles. That is the right kind of self-assessment.\n\nThe soft spots are real but not fatal. First, the practical protocol uses fixed angles α1=π/8, α2=0 because Fig. 5 shows faster convergence, but Porter-Thomas convergence for that ensemble is only checked up to N=100. The advertised target is >100 qubits. The stress-test note is right: the load-bearing premise that fixed-angle circuits are near-Porter-Thomas at the depths used is an extrapolation. It may hold, but the paper does not show it. Second, the noisy validation rests on a single N=64, n=1 setup, with no error bars on Fig. 8. That is thin for a method whose main selling point is practical benchmarking beyond 50 qubits. Third, the fidelity estimator is only tested against fidelity computed from the same noise model, not against hardware. That is acceptable for a first paper, but it limits how strongly one can claim the method will work on real devices.\n\nI would not call this a desk reject. The central estimator is well motivated, the numerics are consistent, and the limitations are stated rather than hidden. What it needs is a follow-up with either a derivation of the noise response, error bars, a hardware demonstration, or at least a scaling analysis for the fixed-angle ensemble. For a reader who wants to extend cross-entropy benchmarking to symmetry-constrained circuits, this is a useful reference. I would send it to peer review; the open questions are exactly what a good referee can help sharpen.","headline":"MLXEB is a plausible, well-tested extension of LXEB to particle-number-conserving circuits, but the fixed-angle Porter-Thomas premise for >100 qubits is an extrapolation, not a proven fact.","tokens_in":14728,"tokens_out":633,"would_cite":true,"duration_ms":8208,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P68"],"pacs":["03.67.-a","03.67.Lx"],"model":"deepseek-v4-flash","headline":"The paper claims that fidelity benchmarking based on linear cross-entropy can be extended beyond 100 qubits by conserving particle number, with a modified estimator that tracks circuit fidelity under low noise.","keywords":["linear cross-entropy benchmarking","MLXEB","particle-number conservation","U(1)-symmetric random circuits","XXZ gates","fidelity estimation","Porter-Thomas distribution","quantum benchmarking"],"falsifier":"A direct way to test the central claim is to simulate the fixed-angle ensemble at N=144 or 196, n=3, at depths Nd near the fitted scaling value, and compare the sorted probability histogram to Dn $e^{{-Dn p}}$; if the distribution's deviation from Porter-Thomas remains visible (for example, a KL divergence that does not decrease with N) or if FMLXEB,n at p_noise=1e-3 disagrees with an independently computed gate-level fidelity beyond the sampling error, the normalization behind MLXEB fails.","tokens_in":13649,"feed_emoji":"⚛️","tokens_out":6217,"duration_ms":53715,"temperature":0.7,"pith_summary":"This paper argues that linear cross-entropy benchmarking (LXEB), the standard fidelity estimator for random-circuit experiments, can be extended to devices beyond 100 qubits by restricting the random circuits to conserve particle number. The proposed estimator, MLXEB, evaluates the cross entropy between measured bitstrings and the ideal distribution computed inside the fixed particle-number-n subspace, whose dimension Dn = C(N,n) is polynomially manageable when n is small. The paper claims that, once the ideal output distribution has converged to the rescaled Porter-Thomas distribution on that subspace and noise is weak, FMLXEB,n tracks the true circuit fidelity. A sympathetic reader would care because this is a path to benchmarking large circuits containing non-Clifford gates that ordinary LXEB cannot classically simulate. The authors support the claim with noiseless simulations up to N=196 qubits, noisy simulations on 64 qubits at p_noise ~ 1e-3, and a reported noiseless simulation feasibility up to 1000 qubits.","feed_headline":"Symmetry rule extends quantum benchmarking past 100 qubits","feed_subtitle":"Keeping the particle number fixed keeps the ideal output classically computable, so fidelity estimates still hold at large scale.","key_machinery":"The load-bearing object is the particle-number-n subspace of dimension Dn = C(N,n), together with the rescaled Porter-Thomas formula Pn(pn)=Dn $e^{{-Dn pn}}$ that is assumed to describe the ideal output probabilities inside that subspace. The circuits are layers of U(1)-symmetric XXZ gates w(alpha1,alpha2)=exp(i[alpha1(X X+Y Y)+alpha2 Z Z]) on a square lattice, with either random angles or fixed angles alpha1=pi/8, alpha2=0; fixing the angles and placing the initial particle near the lattice center accelerates convergence to Porter-Thomas. The estimator FMLXEB,n in Eq. (11) renormalizes the standard LXEB formula by ns/Ns and Dn/ns, so that noiseless sampling gives 1 and fully depolarized sampling gives 0. The authors also introduce a finite-size scaling f(Nd)=((FMLXEB,n-1)/(Dn-2))^{1/(n+a)} with a=0.3, whose stretched-exponential fit predicts the circuit depth needed for reliable benchmarking at larger system sizes.","core_discovery":"The paper's central claim is that fidelity estimation by cross entropy survives a symmetry restriction that makes the ideal distribution classically computable. Defining FMLXEB,n in Eq. (11) as (ns/Ns)[(Dn/ns) Σ pn(x'_j) - 1], with ns the measurement outcomes in the n-particle subspace and pn the noiseless probabilities there, the authors maintain that FMLXEB,n coincides with the circuit fidelity F when the output distribution approaches the rescaled Porter-Thomas distribution Pn(p)=Dn $e^{{-Dn p}}$ and the depolarizing noise rate is low. They demonstrate numerically that square-lattice circuits of U(1)-symmetric XXZ gates, with either random or fixed gate angles, produce this distribution for N=36 to 196 qubits, and that FMLXEB,n converges to 1 noiselessly and matches the true fidelity at p_noise around 1e-3. The method thereby claims access to benchmark regimes beyond the roughly 50-qubit classical simulation wall of ordinary LXEB while remaining compatible with non-Clifford random circuits.","pith_inferences":["If the fixed-angle U(1)-symmetric ensemble is later proven to be a unitary t-design (meaning its statistical moments match Haar-random unitaries), the numerical convergence seen here would become a theorem, and the method would need no per-circuit random angle sampling at all.","The same subspace-truncation idea used for noisy simulation, retaining particle sectors above n, could be adapted to extract noise rates or to benchmark hardware with U(1) charge conservation, such as fermionic simulators.","A natural next test is to apply MLXEB to N>200 at n=2 or 3 with the fixed-angle circuits; if the random-angle scaling relation does not reappear, depth planning for those circuits will need a separate heuristic.","If per-gate error rates continue to drop toward 1e-4, the usable depth and system size for accurate MLXEB estimates would grow roughly logarithmically with inverse noise, potentially reaching the 1000-qubit simulation capability the authors report."],"forward_implications":["MLXEB should benchmark the fidelity of U(1)-symmetric circuits on devices with more than 100 qubits whenever n=O(1), a scale beyond ordinary LXEB's exact-simulation reach.","For 64-qubit circuits run at gate noise around 1e-3, MLXEB with n=1 estimates the overall fidelity at depth Nd approximately 30 within statistical uncertainty.","In the random-angle ensemble, the depth dependence obeys a single scaling curve, so the required depth for a target fidelity can be predicted for larger devices before running them.","Fixed-angle circuits converge to the rescaled Porter-Thomas distribution with less depth, which is useful on hardware with a limited coherent gate budget; however, no scaling relation is yet known for that ensemble.","Because the circuits use non-Clifford XXZ gates, the method covers benchmarking scenarios that Clifford-based benchmarking methods cannot address."],"supporting_citations":[{"why":"Defines standard LXEB and the fidelity estimator that MLXEB modifies.","marker":"[2]"},{"why":"Introduces random circuit sampling and the statistical framework for cross-entropy benchmarking.","marker":"[26]"},{"why":"Supplies the Porter-Thomas distribution and the statistical behavior behind LXEB fidelity.","marker":"[27]"},{"why":"Shows that particle-number-conserving random circuits converge to Porter-Thomas within the fixed-n subspace.","marker":"[38]"},{"why":"Provides the convergence and unitary t-design properties of U(1)-symmetric random circuits used as the basis for MLXEB.","marker":"[39]"},{"why":"Reports that fixed two-qubit gate angles enhance global randomness, motivating the fixed-angle configuration tested here.","marker":"[42]"},{"why":"Supplies the Monte Carlo wavefunction method used for noisy state-vector simulations.","marker":"[45]"},{"why":"Provides the custom simulator routines used to access full probability distributions for large systems.","marker":"[46]"}],"fun_headline_variants":["Symmetry trick benchmarks quantum chips past 100 qubits","Particle-number rule extends quantum benchmarking beyond 100 qubits","MLXEB: fidelity testing for large quantum devices","New cross-entropy variant scales quantum testing past 100 qubits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The fragile premise is that the fixed-angle XXZ circuits (alpha1=pi/8, alpha2=0) on a square lattice produce the rescaled Porter-Thomas distribution over the whole n-particle subspace at the depths used; the paper verifies this numerically only up to N=100 for fixed angles (and N=196 for random angles), and it leaves open which t-design property the fixed-angle ensemble satisfies.","fun_headline_variants_meta":{"raw":{"variants":["Symmetry trick benchmarks quantum chips past 100 qubits","Particle-number rule extends quantum benchmarking beyond 100 qubits","MLXEB: fidelity testing for large quantum devices","New cross-entropy variant scales quantum testing past 100 qubits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000161,"raw_usage":{"total_tokens":1240,"prompt_tokens":956,"completion_tokens":284,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":572,"completion_tokens_details":{"reasoning_tokens":215}},"tokens_in":572,"tokens_out":284,"duration_ms":3507,"temperature":1.0,"reasoning_tokens":215,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:03:31.209585+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct way to test the central claim is to simulate the fixed-angle ensemble at N=144 or 196, n=3, at depths Nd near the fitted scaling value, and compare the sorted probability histogram to Dn $e^{{-Dn p}}$; if the distribution's deviation from Porter-Thomas remains visible (for example, a KL divergence that does not decrease with N) or if FMLXEB,n at p_noise=1e-3 disagrees with an independently computed gate-level fidelity beyond the sampling error, the normalization behind MLXEB fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines standard LXEB and the fidelity estimator that MLXEB modifies."},{"cited_title":"Harper, W","cited_arxiv_id":null,"evidence_quote":"Introduces random circuit sampling and the statistical framework for cross-entropy benchmarking."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Porter-Thomas distribution and the statistical behavior behind LXEB fidelity."},{"cited_title":"A Bottom-up Approach to Constructing Symmetric Variational Quantum Circuits","cited_arxiv_id":"2308.08912","evidence_quote":"Shows that particle-number-conserving random circuits converge to Porter-Thomas within the fixed-n subspace."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the convergence and unitary t-design properties of U(1)-symmetric random circuits used as the basis for MLXEB."},{"cited_title":"Mitsuhashi, R","cited_arxiv_id":null,"evidence_quote":"Reports that fixed two-qubit gate angles enhance global randomness, motivating the fixed-angle configuration tested here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the custom simulator routines used to access full probability distributions for large systems."}],"review_version":1}