{"id":"1d82ef9a-adf8-408c-9494-8260e425d67a","arxiv_id":"2505.22390","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"CAB benchmarks a 44-qubit parallel CZ gate at 63.09% and a 52-qubit version at 38.92%, and global-fidelity optimization improves a 6-qubit parallel CZ gate from about 88.7% to 92.04%.","lead":"This paper reports benchmarking of multi-qubit gate fidelities on up to 52 qubits in a superconducting processor, using the character-average benchmarking protocol. The authors show that optimizing gates using the global fidelity of a parallel gate outperforms tuning each gate individually.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 52-qubit CAB fidelity rests on an unverified non-negativity assumption: λ and −λ are indistinguishable in the Aλ^{2m} fit, and no quality-parameter evidence is shown for the blue pattern.","rationale":"The sign ambiguity is not a minor technical caveat: it is the difference between reporting the process fidelity and reporting the average absolute value of noisy Pauli eigenvalues. The paper's own limitation statement is explicit, and the missing blue-pattern quality-parameter distribution is precisely the evidence needed to close the gap. This is the weakest point in the strongest claim (52-qubit benchmarking), more central than the post-selected optimization window, which affects a secondary application. Independent support for CAB on single CZ gates (Supplemental Fig. S1) does not cover the large-gate, low-fidelity regime: at 17% or 39% global fidelity, even a small coherent component can make quality parameters negative. The proposed check would settle whether the assumed depolarizing regime actually holds for the 52-qubit blue pattern, and it leaves the reader's conditional verdict unchanged while making the condition explicit.","tokens_in":30448,"tokens_out":8845,"duration_ms":113818,"concrete_test":"Recompute the 52-qubit blue-pattern pure fidelity with sign-corrected quality parameters: take the authors' noise model (Eq. 15, with V=e^{-iΣγ_{kl}Z_{i_k}Z_{i_l}} preceding local depolarizing channels), use independently calibrated p_i≈0.966 (from XEB, Table I) and the measured ZZ-coupling strengths γ_{kl} for the blue-pattern qubit pairs, compute the predicted sign of each sampled λ_w, flip the fitted |λ_w| accordingly, and propagate through Eq. (3) to a corrected pure fidelity. If the corrected value shifts from 38.92% by more than the quoted 0.76% standard error, the non-depolarizing assumption is violated and the headline 52-qubit claim is biased; if it shifts by less, the positive-λ assumption is adequate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that CAB yields credible fidelity estimates up to 52 qubits depends on an assumption the authors themselves flag: in Box 1, step 6, each quality parameter is obtained by fitting f_i(m)=Aλ_i^{2m}, so only even powers of λ_i appear and the data determine |λ_i|, not λ_i. The reported fidelity is the average of these fitted values (Eq. 4), so if any true λ_i is negative, replacing it by |λ_i| biases the fidelity upward. The text explicitly says that for low fidelity and non-depolarizing (unitary) noise, quality parameters may be below 0, and a negative λ is indistinguishable from −λ in the exponential fit. For the orange-pattern 44-qubit gate, violin plots of dressed and local quality parameters support a near-depolarizing positive distribution, although fitted magnitudes alone do not establish sign. For the blue-pattern 52-qubit gate, the paper reports only dressed 31.43%, local 80.75%, and pure 38.92% (Supplemental Table III), with no quality-parameter distribution. At such low global fidelity, even a modest unitary component—for example the ZZ coupling in the authors' own model, Eq. (15), with γ ≳ π/4—can drive individual Pauli eigenvalues negative. The same risk applies to the 46-qubit fully connected gate at 17.42%. Without sign information, the headline 'successfully benchmark gate fidelities up to 52 qubits' is not quantitatively established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports character-average benchmarking (CAB) measurements of large parallel CZ gates and of 'fully connected' gate layers on a 54-qubit superconducting processor. Using a constant number of sampled Z observables, the authors benchmark parallel CZ gates on up to 22 pairs (44 qubits) in one pattern and a 26-pair (52-qubit) gate in a second pattern, with a headline pure fidelity of 63.09%±0.23% for the 44-qubit gate. They define an inter-gate correlation metric from the same CAB data, observe positive short-range correlations in one pattern and some negative correlations in another, and explain the signs with a depolarizing-plus-ZZ-coupling noise model. They also use the global CAB fidelity as a cost function in Nelder-Mead optimization of 3-pair and 2-pair parallel CZ gates, reporting that global-fidelity optimization improves the 6-qubit gate from 87.65% to 92.04% while reducing the correlation from 3.53% to 3.22%, compared with local-fidelity optimization.","tokens_in":30696,"tokens_out":8934,"duration_ms":97012,"significance":"If the results hold, this is a substantial experimental scaling demonstration: CAB with constant sample complexity is used to estimate process fidelities for gates substantially larger than previous cycle-benchmarking demonstrations, and the correlation metric provides a practical crosstalk diagnostic from the same data. The paper includes a direct CAB-versus-cycle-benchmarking validation on three CZ gates, repeated-experiment standard deviations, and an error-propagation analysis of the reported uncertainties; these are genuine strengths. The ZZ-coupling model is simple enough to make falsifiable predictions (positive correlations for two-gate coupling, sign reversal when a third gate couples strongly), and the authors are explicit about the λ sign ambiguity. However, the 52-qubit claim rests on a diagnostic that is not shown for that gate, and the optimization comparison uses a post hoc iteration window; both points need attention before the scaling claim is fully established.","major_comments":[{"comment":"The protocol fits each quality parameter to f_i(m)=Aλ_i^{2m}, so the data determine only |λ_i|. Replacing a negative true λ_i by |λ_i| in Eq. (4) inflates the reported fidelity, and the authors explicitly acknowledge this in the 'Fully connected gate benchmarking' paragraph. That paragraph argues that the near-depolarizing condition is met for the fully connected gate and for the orange-pattern parallel CZ gate, where quality-parameter distributions are shown. For the 52-qubit blue-pattern gate, however, Supplemental Table III reports only dressed/local/pure fidelities and no quality-parameter distribution, so positivity is not established for exactly the gate that supports the 'up to 52 qubits' claim. Given the reported pure fidelity of 38.92%, the authors' own ZZ-coupling model (Eq. (15)) with γ ≳ π/4 can drive individual Pauli eigenvalues negative. The same concern applies to the blue-pattern correlation signs (Fig. 3(d)), since those correlations are computed from the same fidelity estimates. Please add the λ_i distributions for the blue pattern and an independent sign check, or restrict the headline claim to the 44-qubit orange pattern where the diagnostic is provided.","section":"Methods, Box 1 step 6 / Eq. (4); Results, 'Fully connected gate benchmarking'; Supplemental Table III"},{"comment":"The quantitative comparison in the abstract—87.65% to 92.04% and 3.53% to 3.22%—is computed from iterations 100–180, a range that the manuscript states was chosen because it is the 'phase of iterative parameter convergence and stable reference fidelities' (Figure 4 caption; see also Supplement Section II.E). This is a post hoc selection: the same data set was used to identify the stable phase and to estimate the improvement. The conclusion that global-fidelity optimization outperforms local-fidelity optimization should be supported by a pre-specified selection rule (for example, the last K iterations or an explicit convergence criterion applied identically to both runs), or by full optimization curves showing that the comparison is insensitive to the chosen window. As written, the headline optimization gain could be an artifact of the chosen window rather than of the objective function.","section":"'Parallel CZ gate optimization', Figure 4, Supplemental Tables V and VI"},{"comment":"For the 52-qubit blue-pattern gate, the dressed and local twirling fidelities are each estimated from two circuit depths only ({0,1} and {1,2}, respectively). With two points the exponential fit is exact and provides no goodness-of-fit check of the assumed Aλ^{2m} form, nor any estimate of model error. The supplement itself notes that with two survival probabilities 'there will be no fitting error'; this means the reported standard error reflects only measurement propagation, not the validity of the noise model. This is especially important because the 52-qubit gate is in the low-fidelity regime where the noise is not demonstrated to be close to depolarizing. Please include an additional depth, or repeated depth sets, for at least a sub-system or the full 52-qubit gate, and report the resulting stability of the fidelity and of the correlation values.","section":"Supplemental Table III and Supplement Section II.B; 'Parallel CZ gate benchmarking'"}],"minor_comments":[{"comment":"The citation '[28? , 29]' contains an unresolved placeholder and should be replaced with the actual reference.","section":"Introduction"},{"comment":"The name 'fully connected gate' is potentially misleading for a brickwork layer of CZ gates on a ring; please add a one-sentence definition or choose a less suggestive term.","section":"Results, 'Fully connected gate benchmarking'"},{"comment":"Please state precisely how 'distance between CZ gates' is counted when two qubit pairs are not connected by a single edge; the current phrase 'minimal line count' is ambiguous.","section":"Figure 3 caption"},{"comment":"The Hoeffding bound is stated for λ_i, but the experimental estimate obtained from the fit is the magnitude |λ_i|; please clarify whether the inequality applies to the signed quality parameter or to the estimated magnitude.","section":"Methods, Eq. (5) and Box 1 step 6"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is well within the scope of a high-impact quantum-information journal if the 52-qubit claim is properly supported. Please ask the authors for the blue-pattern quality-parameter data and a robustness analysis of the optimization window; both appear to be obtainable from data already taken. The central methodology is otherwise sound, and the paper does not overstate the novelty of CAB itself."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nYou should know that the headline result is real but narrower than advertised. The genuine advance is the scale: the authors run CAB on parallel CZ gates up to 44 qubits (63.09% pure) and a 46-qubit fully connected gate (17.42% dressed), and they push a 26-pair parallel CZ gate to 52 qubits (38.92% pure). That is the largest individual-gate benchmarking I've seen, and the validation of CAB against cycle benchmarking on three CZ gates, with lower variance, is a useful data point. The correlation metric is simple—a normalized deviation between global fidelity and the product of local fidelities—but it demonstrably picks out nearby gates and, in the blue pattern, long-range interactions, which is a nice empirical contribution.\n\nThe optimization comparison is also informative, though the abstract oversells it. Comparing the final iterative fidelities, global-fidelity optimization gives 92.04% and local-fidelity optimization gives 87.65%. But the actual starting point (the reference fidelity) is around 88.7% for both; the local objective didn't improve on its reference, it slightly degraded. So the phrase \"from 87.65% to 92.04%\" in the abstract misstates what happened. The authors do correctly report the iterations-100-180 window as the converged phase, but that window is post hoc. This is a presentational flaw, not a fatal one.\n\nThe real soft spot is the λ-sign ambiguity, which the stress-test rightly flags. In Box 1, the fit is to Aλ^{2m}, so λ and −λ are indistinguishable. The authors state this explicitly and note that low-fidelity, non-depolarizing noise can push quality parameters negative. For the orange pattern and the fully connected gates they show violin plots of quality parameters, but those plots are of fitted magnitudes, not signed values, so they cannot establish positivity. For the 52-qubit blue pattern they show no quality-parameter distribution at all. At 38.92% pure fidelity, a modest unitary component—e.g., the ZZ coupling in their own model with γ ≳ π/4—can flip some eigenvalues negative. If that happens, the reported fidelity is biased upward, possibly substantially. The same risk applies to the 17.42% fully connected gate. This doesn't invalidate the scaling trend or the correlation analysis, but it does mean the headline \"successfully benchmarked up to 52 qubits\" is not quantitatively established for the lowest-fidelity gates.\n\nThe noise model is also more qualitative than the text claims: the γ parameters are taken from known coupling strengths and gate time, not fitted to the correlation data, and the agreement is order-of-magnitude. That's fine as a post hoc explanation, but it isn't a confirmed quantitative model.\n\nBottom line: this deserves a serious referee. The scale alone justifies careful review, the issues are addressable, and the authors are honest about the main limitation. If I were editing, I'd send it out with a request that the authors provide signed quality-parameter distributions (or rigorous bounds) for the low-fidelity gates, and fix the abstract. It would be a useful citation for anyone working on large-gate benchmarking.","headline":"Large-scale CAB benchmarking is real and worth referee time, but the 52-qubit fidelity inherits an unverified λ-sign assumption and the abstract oversells the optimization comparison.","tokens_in":31328,"tokens_out":4359,"would_cite":true,"duration_ms":46659,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper reports the largest quantum-gate fidelity benchmarks to date, reaching 52 qubits, and shows that optimizing a parallel gate with its global fidelity outperforms local optimization.","keywords":["quantum gate benchmarking","character-average benchmarking","parallel CZ gates","crosstalk","gate correlation metric","ZZ coupling noise","superconducting quantum processor","gate optimization"],"falsifier":"Run CAB on the same 52-qubit parallel CZ gate with at least three circuit depths, such as m = 0, 1, 2, and 3, and inspect whether the fitted exponential decays all have positive amplitudes and consistent quality parameters; if any survival-probability curve can be fit equally well by a negative quality parameter, or if the fidelity estimate shifts by more than its error bars when depths are added, the near-depolarizing assumption fails for that gate.","tokens_in":30201,"feed_emoji":"⚛️","tokens_out":7331,"duration_ms":75168,"temperature":0.7,"pith_summary":"The paper claims that a shallow-circuit protocol called character-average benchmarking (CAB) can deliver trustworthy process-fidelity estimates for very large parallel quantum gates, and it reports measurements on a 54-qubit superconducting processor for gates spanning up to 52 qubits. On a 44-qubit parallel controlled-Z (CZ) gate the measured fidelity is 63.09±0.23%, and on a 52-qubit parallel CZ gate the pure fidelity estimate is 38.92%. The authors also introduce a normalized correlation between the global gate fidelity and the product of its local gate fidelities, use it to detect crosstalk among parallel CZ gates, and show that optimizing with the global fidelity outperforms optimizing gate-by-gate: a 6-qubit parallel CZ gate improves from 87.65% to 92.04% while inter-gate correlation drops from 3.53% to 3.22%. If correct, the work offers a practical route to calibrating and optimizing the large multi-qubit gate layers needed in near-term quantum processors and quantum error-correction circuits.","feed_headline":"Fidelity measured for a 52-qubit parallel gate","feed_subtitle":"Shallow-circuit benchmarking also lifts a 6-qubit CZ gate from 87.65% to 92.04% fidelity.","key_machinery":"The central object is the quality parameter of the noise channel under Pauli twirling, estimated by fitting survival probabilities to $Aλ^{{2m}}$ for circuits of depth m. The protocol inserts random local Clifford and Pauli gates around alternating U and $U^{{-1}}$, then averages a constant number of sampled Pauli observables, so the classical postprocessing cost is independent of qubit number. The correlation metric is (F(U) − ∏F(U_i)) / √(F(U)∏F(U_i)), and the physical mechanism used to explain the data is a composite noise channel Λ = Λ_V ∘ ⊗_i Λ_{p_i}, where Λ_{p_i} is depolarizing and Λ_V is unitary ZZ coupling V = exp(−iΣ γ_{kl} Z_{i_k}Z_{i_l}).","core_discovery":"The paper demonstrates that character-average benchmarking can estimate the process fidelity of a Clifford gate on a shallow circuit whose depth does not scale with the gate's order. For a parallel CZ gate made of 22 independent CZ pairs on 44 qubits, it reports a purified fidelity of 63.09±0.23%; for the 26-pair blue pattern on 52 qubits, it reports a pure fidelity of 38.92%. The paper defines an inter-gate correlation as the normalized gap between the global fidelity and the product of local fidelities, detects positive and negative correlations consistent with pairwise ZZ couplings, and shows that using the global fidelity as the optimization target improves a 6-qubit parallel CZ gate from 87.65% to 92.04% while decreasing correlation from 3.53% to 3.22%.","pith_inferences":["Editorial extension: the 52-qubit blue-pattern fidelity should be read as conditional on near-depolarizing noise; the paper shows a quality-parameter distribution only for the orange pattern, so re-measuring the blue pattern with additional circuit depths would test whether all fitted quality parameters remain positive.","Editorial extension: the same correlation-versus-distance scatter analysis could be used to map the crosstalk graph of any qubit array and to decide where to place parallel gates to suppress correlated errors.","Editorial extension: the layer-by-layer decomposition of circuits, clustering strongly correlated gates and optimizing each cluster globally, is a concrete route to scale the method beyond six qubits; the paper suggests the idea but does not demonstrate it."],"forward_implications":["Large multi-qubit gate layers can be benchmarked end-to-end with shallow circuits, bypassing the exploding gate-order problem that blocks cycle benchmarking for fully connected gates.","The global fidelity of a parallel gate is a usable optimization target: on a 6-qubit, three-pair CZ gate it raises fidelity from 87.65% to 92.04% and lowers correlation from 3.53% to 3.22% relative to local-fidelity optimization.","Correlation values computed from the same experiment act as a crosstalk diagnostic: magnitudes track coupling strength, two-gate ZZ coupling gives positive correlation, and adding a strongly coupled third gate can flip the sign of the correlation.","Individual CZ fidelities in the orange pattern remain near 98% as the parallel gate grows, indicating weak short-range crosstalk, which the paper reads as favorable for quantum error correction."],"supporting_citations":[{"why":"Introduces the character-average benchmarking protocol and proves its sample complexity is independent of qubit number, the basis of the paper's scalable method.","marker":"[33]"},{"why":"Provides the interleaved randomized benchmarking equation used to convert dressed CAB fidelities into pure gate fidelities.","marker":"[22]"},{"why":"Describes cycle benchmarking, the comparison baseline the paper uses to validate CAB and to show CAB's smaller statistical fluctuations.","marker":"[25]"},{"why":"Supplies the simplex-based optimization algorithm used for the parallel CZ gate optimization.","marker":"[37]"},{"why":"Supplies the reference for the optimization approach used to tune CZ gate parameters.","marker":"[39]"},{"why":"Describes the design of the superconducting processor platform on which the experiments are run.","marker":"[42]"},{"why":"Describes the all-microwave coupler scheme used to realize the fixed-duration CZ gates.","marker":"[46]"}],"fun_headline_variants":["Shallow-circuit benchmarking scales to 52-qubit gates","Global fidelity optimization boosts 6-qubit CZ gate fidelity","52-qubit gate fidelity measured via character-average benchmarking","Crosstalk quantified in 44-qubit parallel CZ gate","Benchmarking up to 52 qubits with shallow circuits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The extracted fidelity is trustworthy only if the noise on the gate is close to depolarizing, because the fitting procedure cannot tell a negative quality parameter from its positive mirror image.","fun_headline_variants_meta":{"raw":{"variants":["Shallow-circuit benchmarking scales to 52-qubit gates","Global fidelity optimization boosts 6-qubit CZ gate fidelity","52-qubit gate fidelity measured via character-average benchmarking","Crosstalk quantified in 44-qubit parallel CZ gate","Benchmarking up to 52 qubits with shallow circuits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000907,"raw_usage":{"total_tokens":3906,"prompt_tokens":957,"completion_tokens":2949,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":2865}},"tokens_in":573,"tokens_out":2949,"duration_ms":22991,"temperature":1.0,"reasoning_tokens":2865,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:08:46.712332+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run CAB on the same 52-qubit parallel CZ gate with at least three circuit depths, such as m = 0, 1, 2, and 3, and inspect whether the fitted exponential decays all have positive amplitudes and consistent quality parameters; if any survival-probability curve can be fit equally well by a negative quality parameter, or if the fidelity estimate shifts by more than its error bars when depths are added, the near-depolarizing assumption fails for that gate.","supporting_citations":[{"cited_title":"Barends, J","cited_arxiv_id":null,"evidence_quote":"Provides the interleaved randomized benchmarking equation used to convert dressed CAB fidelities into pure gate fidelities."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the all-microwave coupler scheme used to realize the fixed-duration CZ gates."}],"review_version":1}