{"id":"1fb8fcab-659c-4fa0-af1c-304f44c483a5","arxiv_id":"2505.17731","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"On IBM Brisbane, sequential and hybrid circuits discriminate unitary channels more reliably than wide parallel circuits, until depth noise dominates.","lead":"This paper benchmarks how well IBM's Brisbane quantum processor can tell apart two quantum operations when it gets many chances. It finds that using many qubits in parallel hurts more than running the operations one after another, until the circuits get too deep.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Example 1's 'depth' axis is virtual: IBM RZ gates are phase-frame updates, so sequential RZ(π/N) circuits do not scale in physical depth; Figs. 7 and 8 do not test depth tolerance.","rationale":"The paper's central claim is that architectures minimizing entanglement overhead are more resilient to hardware noise as long as depth stays below a threshold. For that claim to hold, the experiments must actually vary physical circuit depth. In Example 1, the unknown channel is RZ, which IBM implements as a virtual phase update; the sequential and hybrid circuits therefore have no physical depth scaling in the black-box segment. This means Fig. 7(a) and Fig. 8 measure only the growth of entangling-gate count in state preparation and measurement, not a depth-versus-entanglement tradeoff. The reader's identified weakness, the post-hoc bit-flip correction in Section 5.5, is real and should be addressed, but it is not the most load-bearing issue for the abstract's central claim: even a fully verified bit-flip correction would leave Example 1 unable to support the depth-resilience statement. Example 2 does provide independent evidence for a depth threshold because its black box contains physical √X gates, so the paper is not beyond repair. The appropriate verdict remains CONDITIONAL: the authors should remove the depth interpretation from Example 1, report repeated-run statistics, and either verify the bit-flip artifact with device calibration or drop corrected data from the central claims. I therefore disagree with the reader's choice of weakest assumption while agreeing with the conditional assessment.","tokens_in":15550,"tokens_out":14660,"duration_ms":121868,"concrete_test":"Use the authors' public GitHub code to transpile the Example 1 circuits for (w,d)=(1,4), (1,8), (1,12) and for a hybrid case such as (w,d)=(4,3) to IBM Brisbane's basis (ecr, sx, x, rz) with optimization_level=3. Report the number of ECR gates and the physical circuit depth, excluding virtual RZ, versus the nominal depth d. If the compiled sequential circuits have zero ECR gates and constant physical depth as N increases, the depth axis in Fig. 7(a) is not physical. As a control, repeat the comparison with the Example 2 black box, which contains √X, and verify that physical depth does increase with d.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.1 introduces Example 1, discriminating identity from RZ(π/N) with no mid-circuit processing. For a purely sequential scheme (width w=1, depth d=N), the black-box part is RZ(π/N) applied N times to one qubit. On IBM Brisbane, RZ is not a physical pulse but a virtual phase update; whether or not consecutive RZ gates are explicitly merged during transpilation, the black-box segment contributes no physical gate error and no physical duration. The same holds for each qubit in the hybrid rectangular schemes: d applications of RZ(π/N) become a single effective rotation RZ(dπ/N). Consequently, Figs. 7(a) and 8, which vary width w while holding w·d=N, vary only the size of the GHZ-preparation and measurement circuits; they do not vary physical circuit depth. The paper's conclusion in Section 5.4 that entangling-gate error dominates over decoherence from gate depth, and the abstract's depth-threshold claim, therefore cannot be inferred from Example 1. Example 2, where the unknown channels contain physical √X gates, does exercise a real depth axis and may support a threshold, but the first experiment does not. The bit-flip correction in Section 5.5 is a separate validity issue; even if fully verified, it does not fix this confound.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports experiments on the IBM Quantum processor Brisbane for the multiple-shot discrimination of two qubit unitary channels, comparing purely parallel, purely sequential, and rectangular sequentially-paralleled schemes. Two examples are studied: distinguishing identity from RZ(π/N) with no mid-circuit processing (Example 1), and distinguishing U=√X RZ(−π/2N)√X from V=√X RZ(π/2N)√X using X and √X as processing gates (Example 2). In theory all N=wd schemes achieve perfect discrimination when the angle condition θ(V†U)=π/N holds; the experiments instead show performance degradation that depends on circuit width and depth. The authors also compare CNOT- and ECR-based transpilation strategies, apply M3 measurement-error mitigation, and introduce a post-hoc label-swapping correction for bit-flip anomalies. The central empirical claim is that architectures minimizing entanglement overhead are more resilient to hardware noise as long as circuit depth does not exceed a threshold.","tokens_in":15792,"tokens_out":7212,"duration_ms":66467,"significance":"If the conclusions were fully supported, this would be a useful experimental contribution to NISQ-era benchmarking of quantum channel discrimination, since it tests a theoretically motivated family of discrimination circuits on real hardware and makes the data openly available. The paper has clear strengths: raw and mitigated results are reported together, several transpilation strategies are compared, runs were repeated on different dates, and the data are deposited on GitHub and Zenodo. However, the significance is currently limited by three issues: the virtual-depth confound in Example 1, the unverified and data-dependent bit-flip correction, and the absence of statistical uncertainty estimates for the main quantitative claims. These issues affect the abstract's and conclusion's central statements, so the empirical conclusions should be regarded as preliminary until the concerns are addressed.","major_comments":[{"comment":"In Example 1, the 'depth' axis is virtual. On IBM Brisbane, RZ gates are implemented as frame updates, so d successive RZ(π/N) applications on the same qubit compile to a single virtual rotation RZ(dπ/N) with no additional physical duration or gate error. Consequently, both the purely sequential scheme in Fig. 7(a) and the rectangular schemes in Fig. 8 vary only the width of the GHZ preparation and measurement circuits when w·d=N is held fixed; they do not vary the physical depth of the unknown-channel segment. The statement in §5.4 that the results show that the primary source of performance degradation is multi-qubit gate error rather than 'decoherence from circuit depth alone' is therefore not supported by Example 1. The abstract's depth-threshold claim should be based on Example 2, whose unknown channels contain physical √X gates and thus provide a genuine depth axis, or on additional experiments that introduce physical depth without entangling gates.","section":"§5.1, §5.4, Figs. 7 and 8"},{"comment":"The label-swapping correction is applied post hoc whenever the raw success probability drops below 0.5. Because the theoretical prediction is p_succ=1, this procedure guarantees that the corrected value lies above 0.5 and biases the data toward the theoretical expectation. The manuscript itself describes the underlying bit-flip artifact as a hypothesis requiring further investigation, and the cited evidence is not sufficient: the fact that M3 error mitigation has no effect is expected if the error is not a measurement-assignment error, and it does not establish that the flips are global, systematic, or independent of the prepared state. I ask the authors to verify the artifact with dedicated calibration experiments (for example, GHZ states of variable width with known output parity), to state an a priori rule for when swapping is permissible, and to report both raw and corrected values throughout. As written, the corrected probabilities in Fig. 9 and the 90% per-qubit accuracy used in §6.3 rest on an unverified assumption.","section":"§5.5, Fig. 9"},{"comment":"The quantitative comparison of schemes lacks error bars and uncertainty estimates. Each circuit uses 10,000 shots, so binomial sampling error is small but nonzero, and the runs come from a single device over different dates with no explicit treatment of calibration drift or correlated errors. Without confidence intervals or repeated measurements, statements such as 'we received p_succ=0.56765 that is better than any optimal scheme' (§6.3) and the threshold behavior claimed in §6.2 are not yet established. The authors should provide per-point confidence intervals, repeat the key comparisons across device calibrations, and state whether the qualitative trends are stable under those repetitions.","section":"§6.2, §6.3, Figs. 10 and 11"}],"minor_comments":[{"comment":"The text states 'for k<w−k we guess Φ=ΦU'; the second guess should be Φ=ΦV. Please also clarify which figure supports the 'around 90% accuracy for each qubit' used for the N=1024 suboptimal protocol, since Fig. 10 shows only N=4, 16, and 32.","section":"§6.3"},{"comment":"The displayed θ expression after Eq. (16) contains the same RZ(−π/(2N)) factor on both sides; it should be V*†U*, with the opposite-sign RZ angle or an explicit dagger, in order to yield θ(RZ(π/N)^d)=π.","section":"§6.1, text after Eq. (16)"},{"comment":"The introduction of N=w·d mentions 'where k and l are natural numbers'; this should refer to w and d.","section":"§4.3"},{"comment":"There are minor language and typographical issues: 'does not overpass threshold value' should read 'does not exceed threshold value', and 'dimentions' should be 'dimensions'. The figure labels 'Numberofshots' also lack spaces.","section":"Abstract and §5.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a single-device experimental benchmark. Its topic fits a quantum-information experiment or benchmarking venue, but the title's 'IBM Q computers' overstates the device coverage, and the abstract's claim is stronger than what the data currently support. The main obstacles for publication are the virtual-depth confound in Example 1, the post-hoc bit-flip correction, and the lack of statistical analysis; all three appear to be addressable within the scope of a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper contains a real experimental comparison nobody else has run: hybrid w-by-d rectangular schemes for unitary channel discrimination on IBM Brisbane, plus a careful transpilation study and a suboptimal majority-voting experiment. Data and code are public. Second, the main depth-versus-width conclusion is weaker than the abstract suggests, because Example 1 uses RZ gates that are virtual phase updates on IBM hardware. The 'depth' axis in Figs. 7 and 8 does not actually add physical gates or decoherence; those figures vary only the GHZ preparation and measurement width. So the claim that entangling errors dominate over depth noise is not supported by Example 1. Example 2, which uses physical √X gates, does exercise a real depth axis and does show a threshold—that is the paper's strongest evidence, but on one device and without error bars.\n\nThe qualitative trend in the raw data is visible and worth taking seriously. What is genuinely new: the rectangular scheme comparison, the 11-qubit transpilation results showing roughly a 20% accuracy gain with topology-aware fixed mapping for XOR measurement, and the observation that a suboptimal 32-qubit voting strategy at N=1024 beats all optimal schemes. Those are useful design data points for NISQ benchmarking and phase-estimation-style black-box tasks.\n\nThe soft spots are real, though not equally soft. The post-hoc label swapping in Section 5.5 is the most serious: the authors swap the expected answer sets whenever p_succ drops below 0.5, then use the corrected numbers in figures and in the 90% per-qubit accuracy that predicts the voting strategy. They hypothesize a systematic bit-flip artifact but never verify it with calibration data or a device-level model; M3 mitigation having no effect is not evidence for a global flip. This correction should either be substantiated or removed from the central claims. The lack of error bars is a minor issue given 10,000 shots per circuit, but repeated runs would help.\n\nWho this is for: experimentalists and benchmark designers. It deserves a serious referee, but the referee should demand the RZ confound be acknowledged and the label swapping be justified. If the authors verify the artifact or drop the corrected data and add uncertainty quantification, it would be a solid benchmarking paper.","headline":"Useful NISQ benchmarking data undermined by a virtual-RZ depth confound in Example 1 and an unverified label-swapping correction.","tokens_in":16385,"tokens_out":2822,"would_cite":false,"duration_ms":20959,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["03.67.-a"],"model":"deepseek-v4-flash","headline":"On the IBM Brisbane processor, the dominant error source in multiple-shot unitary-channel discrimination is multi-qubit entangling-gate error, not circuit depth alone.","keywords":["quantum channel discrimination","unitary channels","multiple-shot discrimination","parallel and sequential schemes","hybrid schemes","entangling gate errors","noise resilience","IBM Brisbane processor"],"falsifier":"Run the same five-plus-qubit discrimination circuits immediately after calibration, randomizing the assignment of the two answer sets across otherwise identical runs and testing two different logical-to-physical mappings. If the 'global bit flip' appears only for one assignment, or follows the logical labels rather than the physical qubits, the systematic-artifact hypothesis is refuted and the corrected probabilities in the figures would need re-baselining.","tokens_in":15295,"feed_emoji":"⚛️","tokens_out":8913,"duration_ms":62524,"temperature":0.7,"pith_summary":"This paper asks whether the theoretical promise of multiple-shot quantum channel discrimination survives contact with a real noisy processor. For two unitary channels whose single-copy overlap is too small to distinguish, theory says that parallel, sequential, and rectangular hybrid schemes all become perfect once the number of copies $N$ satisfies $N\\theta(V^\\dagger U)\\geq\\pi$, where $\\theta$ is the arc length of the spectrum of $V^\\dagger U$. Running these circuits on the IBM Brisbane processor, the authors find that neither pure parallel nor pure sequential circuits perform well: parallel circuits drown in entangling-gate errors as the GHZ discriminator widens, while very deep sequential circuits lose to decoherence. The paper's central claim is that hybrid sequentially-paralleled circuits, which minimize entanglement overhead while keeping depth below a threshold, are the most resilient in practice, and that suboptimal majority-voting strategies can beat theoretically optimal circuits in the heaviest noise regime.","feed_headline":"Entangling gates, not depth, dominate quantum-discrimination errors","feed_subtitle":"Fewer entangling gates, not fewer steps, decides which discrimination circuit wins on today's chips.","key_machinery":"The machine at work is the arc function $\\theta(V^\\dagger U)$, the length of the smallest arc on the unit circle that contains all eigenvalues of $V^\\dagger U$; perfect single-shot discrimination holds iff $\\theta\\geq\\pi$, and $N$ copies give perfect discrimination iff $N\\theta\\geq\\pi$. The paper builds rectangular 'sequentially-paralleled' schemes with width $w$ and depth $d$, $N=wd$, placing $N$ copies of the unknown channel as $d$ layers of $w$ parallel applications. The discriminator is a GHZ-type state produced by a cascade of CNOT or ECR entangling gates, and the measurement is either a shallow 'short' circuit or a deeper XOR-based circuit whose parity bit identifies the channel. The role of this machinery is to make all three schemes theoretically equivalent, all giving $p_{\\mathrm{succ}}=1$, so that any observed difference is attributable to hardware noise rather than to the discrimination strategy itself.","core_discovery":"The empirical discovery is that on the IBM Brisbane processor the dominant error source in multiple-shot unitary-channel discrimination is the multi-qubit entangling gate, not circuit depth alone. In the first example (identity versus $R_Z(\\pi/N)$), purely sequential circuits stay near $p_{\\mathrm{succ}}\\approx 0.96$ for $N$ up to 12, while purely parallel circuits drop from near 1 to below 0.5 as width grows, and hybrid schemes degrade as more entangling gates are added. In the second example ($U=\\sqrt{X}R_Z(-\\pi/2N)\\sqrt{X}$ versus $V=\\sqrt{X}R_Z(\\pi/2N)\\sqrt{X}$), sequential schemes win for small $N$, sequentially-paralleled schemes win for $N=64$ and $N=96$, and at $N=1024$ all optimal schemes fail while an explicitly suboptimal scheme with 32 independent sequential chains and majority voting reaches $p_{\\mathrm{succ}}=0.56765$. The authors conclude that circuit architectures minimizing entanglement overhead while preserving discrimination power are significantly more resilient to hardware noise, provided their depth does not exceed a threshold.","pith_inferences":["A direct test of the paper's global-bit-flip hypothesis would randomize the logical-to-physical qubit mapping across runs; if the flip follows the logical answer sets rather than the physical qubits, the correction is suspect.","The qualitative ranking of schemes on Brisbane may not transfer to devices with different native gate sets or error profiles; the transferable quantity is the per-layer entangling-gate error budget, not the absolute threshold depth.","A natural follow-up measures the same three scheme classes across calibration epochs with varying two-qubit gate error rates, to check whether entangling-gate error is the causal driver rather than crosstalk or measurement error.","Without a noise model for the observed global bit flips, the 90 percent per-qubit accuracy used to predict the suboptimal strategy's performance should be read as an upper bound, not a calibrated estimate."],"forward_implications":["Circuit designers facing noisy hardware should prefer deeper, narrow circuits over wide, shallow ones, because entangling-gate count rather than depth alone drives the error rate.","Rectangular hybrid schemes with intermediate width are the practical operating point for many-copy tasks: wide enough to cut depth, narrow enough to limit entanglement overhead.","Theoretically suboptimal strategies, such as independent sequential runs per qubit with majority voting, can outperform every optimal scheme in heavily noisy regimes such as $N=1024$.","Hardware-aware compilation, using topology-aware ECR circuits with fixed qubit mapping, can recover roughly 20 percent accuracy on 11-qubit XOR-measurement circuits compared with generic CNOT transpilation.","Black-box tasks with many oracle calls, such as quantum phase estimation, should be re-examined under the same depth-versus-entanglement trade-off."],"supporting_citations":[{"why":"Supplies the arc-function scaling condition $N\\theta\\geq\\pi$ for perfect multi-shot discrimination of unitary channels, the theoretical baseline all experiments are designed to meet.","marker":"[44]"},{"why":"Proves entanglement is not necessary for perfect discrimination between unitary operations and gives the condition used for sequential schemes.","marker":"[32]"},{"why":"Provides the theoretical optimality of the parallel scheme that the experiments implement and challenge on hardware.","marker":"[30]"},{"why":"Earlier experimental demonstration of distinguishing unitary gates on an IBM processor, the direct predecessor this work extends to multiple shots and hybrid layouts.","marker":"[38]"},{"why":"Gives the Helstrom bound underlying the probability of correctly guessing between two quantum states after the optimal measurement.","marker":"[5]"},{"why":"Together with Helstrom, yields the formula $p_{\\mathrm{succ}}=1/2+\\frac{1}{4}\\|\\Phi_0-\\Phi_1\\|_\\diamond$ used to connect single-shot channel discrimination to the diamond norm.","marker":"[41]"},{"why":"Provides the M3 measurement-error mitigation technique whose lack of effect is cited as evidence that the observed global bit flips are not ordinary measurement errors.","marker":"[49]"},{"why":"Documents the Eagle R3 processor architecture whose native ECR gates motivate the hardware-specific discriminator and measurement circuits.","marker":"[47]"}],"fun_headline_variants":["Entangling gates, not depth, decide quantum-discrimination wins","Quantum discrimination loses when entanglement overhead grows","On IBM Q, entangling gates outweigh depth in discrimination errors","Minimize entangling gates to win quantum channel discrimination","IBM Brisbane: entanglement overhead, not depth, sinks discrimination"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the global bit-flip pattern seen on circuits with five or more qubits is a systematic device artifact, so swapping the expected answer sets is a valid correction; if the flips are state-dependent or sporadic, the corrected success probabilities are not trustworthy.","fun_headline_variants_meta":{"raw":{"variants":["Entangling gates, not depth, decide quantum-discrimination wins","Quantum discrimination loses when entanglement overhead grows","On IBM Q, entangling gates outweigh depth in discrimination errors","Minimize entangling gates to win quantum channel discrimination","IBM Brisbane: entanglement overhead, not depth, sinks discrimination"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001167,"raw_usage":{"total_tokens":4793,"prompt_tokens":872,"completion_tokens":3921,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":488,"completion_tokens_details":{"reasoning_tokens":3843}},"tokens_in":488,"tokens_out":3921,"duration_ms":20847,"temperature":1.0,"reasoning_tokens":3843,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:41:57.809162+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same five-plus-qubit discrimination circuits immediately after calibration, randomizing the assignment of the two answer sets across otherwise identical runs and testing two different logical-to-physical mappings. If the 'global bit flip' appears only for one assignment, or follows the logical labels rather than the physical qubits, the systematic-artifact hypothesis is refuted and the corrected probabilities in the figures would need re-baselining.","supporting_citations":[{"cited_title":"Statistical distinguishability between unitary operations.Physical Review Letters, 87(17):177901, 2001","cited_arxiv_id":null,"evidence_quote":"Supplies the arc-function scaling condition $N\\theta\\geq\\pi$ for perfect multi-shot discrimination of unitary channels, the theoretical baseline all experiments are designed to meet."},{"cited_title":"Entanglement is not necessary for perfect discrimi- nation between unitary operations.Physical Review Letters, 98(10):100503, 2007","cited_arxiv_id":null,"evidence_quote":"Proves entanglement is not necessary for perfect discrimination between unitary operations and gives the condition used for sequential schemes."},{"cited_title":"Parallel distinguishability of quantum op- erations","cited_arxiv_id":null,"evidence_quote":"Provides the theoretical optimality of the parallel scheme that the experiments implement and challenge on hardware."},{"cited_title":"Distinguishing unitary gates on the ibm quantum processor","cited_arxiv_id":null,"evidence_quote":"Earlier experimental demonstration of distinguishing unitary gates on an IBM processor, the direct predecessor this work extends to multiple shots and hybrid layouts."},{"cited_title":"Quantum detection and estimation theory.Journal of Statistical Physics, 1:231–252, 1969","cited_arxiv_id":null,"evidence_quote":"Gives the Helstrom bound underlying the probability of correctly guessing between two quantum states after the optimal measurement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Together with Helstrom, yields the formula $p_{\\mathrm{succ}}=1/2+\\frac{1}{4}\\|\\Phi_0-\\Phi_1\\|_\\diamond$ used to connect single-shot channel discrimination to the diamond norm."},{"cited_title":"Scalable mitigation of measurement errors on quantum computers.PRX Quantum, 2(4):040326, 2021","cited_arxiv_id":null,"evidence_quote":"Provides the M3 measurement-error mitigation technique whose lack of effect is cited as evidence that the observed global bit flips are not ordinary measurement errors."},{"cited_title":"Accessed on 2025- 05-18","cited_arxiv_id":null,"evidence_quote":"Documents the Eagle R3 processor architecture whose native ECR gates motivate the hardware-specific discriminator and measurement circuits."}],"review_version":1}