{"id":"c553a588-09d2-4514-a555-ce91dae369f2","arxiv_id":"1909.01282","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A randomized-measurement protocol estimates the overlap of two quantum states prepared on separate platforms, with a 10-qubit trapped-ion proof of principle.","lead":"Randomized measurements on two quantum devices can estimate how similar the devices' output states are, without needing to know the states or do full tomography. The paper demonstrates the idea on a 10-qubit trapped ion simulator, comparing measured data with theory.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Appendix E's assertion that unitary errors 'do not lead to false positives' is contradicted by its own Eq. (E6): for orthogonal states with overlapping marginals the estimated overlap is positive at order η².","rationale":"The central derivation of Eq. (2) is sound; I checked the Weingarten calculation in Appendix A. The load-bearing weakness is the robustness analysis. The paper claims imperfections only decrease fidelity and never cause false positives (main text and Sec. E2). However, Eq. (E6) shows the estimated overlap contains a positive term proportional to Σ_k Tr[Tr{k}[ρ1]Tr{k}[ρ2]] that does not vanish when Tr[ρ1ρ2]=0. For any pair of orthogonal states with overlapping single-particle marginals (e.g., the N=2 example), the estimated Fmax is positive of order η², so false positives occur. This matters because cross-platform verification must guard against two different states appearing similar. The paper's numerical checks (Fig. E.1) only test identical states and thus miss this. The fix would be to derive a two-sided error bound or to include calibration of the unitary errors; absent that, the practical claim is conditional. The reader's condition on an actual two-platform demonstration remains valid, and I add the condition that the false-positive issue be resolved.","tokens_in":96117,"tokens_out":15071,"duration_ms":143064,"concrete_test":"Simulate the protocol with N=2 qubits, ρ1=|00⟩⟨00|, ρ2=(|01⟩+|10⟩)(⟨01|+⟨10|)/2, using the Appendix E error model with η1=0, η2=0.1, no projection noise, and many random unitaries U (NU≈500). Compute the estimated Fmax from Eq. (2). If it exceeds the true value 0 by an amount consistent with η², the 'no false positives' claim is falsified. An analytic cross-check: evaluate Eq. (E6) for this pair; the result is E1,2=η²>0.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's robustness claim is one-sided. Appendix E proves only that for ρ1=ρ2 (unit fidelity) unitary errors and depolarization lower the estimated fidelity. For general pairs, Eq. (E6) gives E1,2 = (1 − 2(η1²+η2²)N)Tr[ρ1ρ2] + (η1²+η2²) Σ_k Tr[Tr{k}[ρ1]Tr{k}[ρ2]]. Take N=2, ρ1=|00⟩⟨00|, ρ2=(|01⟩+|10⟩)(⟨01|+⟨10|)/2. Then Tr[ρ1ρ2]=0, but each single-qubit marginal overlap is 1/2, so E1,2≈η²>0. With unit purities, the estimated Fmax is ≈η² while the true fidelity is zero. Thus unitary mismatch can create false positives—exactly the dangerous failure mode for cross-platform verification. The numerical check in Fig. E.1(a) only tests ρ1=ρ2 and cannot see this. The ideal derivation of Eq. (2) is sound, but the practical guarantee that imperfections are conservative is not established and is in fact false for states with overlapping marginals.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a protocol to estimate the overlap Tr[\\rho_1\\rho_2] and purities of two quantum states prepared on separate platforms, using only local randomized measurements with the same local random unitaries communicated classically. From these quantities it defines the mixed-state fidelity Fmax and demonstrates, using data from Ref. [28], experiment-theory fidelities of 10-qubit trapped-ion states and 'experiment-experiment' fidelities obtained by splitting one experimental data set into two halves. Appendix A derives Eq. (2) using Weingarten calculus for product local 2-designs, Appendices C and D study statistical errors and resampling, and Appendix E models the effect of unitary errors and depolarization on the estimators.","tokens_in":96386,"tokens_out":1478,"duration_ms":20955,"significance":"If the protocol's practical guarantees hold, this is a valuable and timely tool: it offers a concrete, scalable-in-practice route to compare unknown states across devices or against a classical target without tomographic overhead, and the paper provides an explicit machine-checkable-style derivation of the central identity, careful resource scaling numerics, and a genuine 10-qubit proof-of-principle from existing trapped-ion data. The paper also correctly identifies and experimentally probes unitary calibration errors, which is the main practical threat to cross-platform operation. The central Eq. (2) is derived, not fitted, and the statistical-error analysis with bootstrap resampling is a useful contribution in itself.","major_comments":[{"comment":"The robustness claim in the main text and Appendix E that unitary errors 'do not lead to false positives' is not established for general pairs of states. For orthogonal states with overlapping marginals, Eq. (E6) gives E_{1,2} ≈ (η_1^2 + η_2^2) Σ_k Tr[Tr_{(k)}[ρ_1] Tr_{(k)}[ρ_2]], which is positive even when Tr[ρ_1 ρ_2] = 0. With unit purities, the estimated Fmax is then ≈ η^2 while the true fidelity is zero. The numerical check in Fig. E.1(a) only considers ρ_1 = ρ_2 and therefore cannot detect this false-positive mechanism. This is load-bearing because the protocol's practical appeal is exactly that imperfections degrade, rather than inflate, the measured fidelity; as written, the claim is one-sided and the dangerous failure mode (unitary mismatch creating spurious agreement) is not excluded.","section":"Appendix E, Eq. (E6)"},{"comment":"The 'experiment-experiment' demonstration does not compare two physical platforms. The data are split into two halves E1 and E2 from the same experimental run of Ref. [28]; both halves estimate the same underlying state, so high Fmax is largely forced by construction and does not test the cross-platform scenario advertised in the title and abstract. This should be clearly labeled as a self-consistency check, and the paper should state explicitly that no two-device demonstration is provided.","section":"Section 'Fidelity estimation with trapped ions', Fig. 4(a,b)"},{"comment":"Fig. E.2 shows that implementing a single random unitary twice (two concatenated unitaries) lowers the estimated fidelity of a known product state from near one to significantly below one, and the paper itself attributes this to unitary errors. In a cross-platform setting, the same random U_A is implemented by two different devices with independent calibration, so the relevant error is the mismatch between U_A^{(1)} and U_A^{(2)}, not a common error. The paper does not quantify this differential error; without a model or a two-device test, the protocol's applicability to genuinely different platforms remains unproven, and the false-positive mechanism of the first major comment becomes the relevant worst case.","section":"Appendix E, Sec. 3 and Fig. E.2"},{"comment":"The central derivation is sound: averaging the local 2-design over each qudit and using the swap trick yields Tr[ρ_{i,A_i} ρ_{j,A_j}] exactly, with no fitting parameters. I stress this as a strength rather than a weakness. The concern is only that the experimental and robustness sections claim a guarantee that the derivation itself does not provide; the derivation is for ideal unitaries, while Appendix E's error analysis is incomplete as described above.","section":"Eq. (2) and Appendix A"}],"minor_comments":[{"comment":"The title and abstract promise 'cross-platform verification', but no two physical platforms are compared anywhere in the paper. Consider rewording to 'cross-platform protocol with single-platform demonstration' or similar, to avoid overclaiming.","section":"Abstract and title"},{"comment":"The definition of Fmax uses 'max{Tr[ρ_1^2], Tr[ρ_2^2]}' but the text in the paragraph and elsewhere sometimes writes Tr[ρ_i^2] with indices interchanged; this is not a technical error, but the notation would be clearer if the two reduced-state labels were consistently distinguished as ρ_{1,A_i} and ρ_{2,A_j} throughout.","section":"Main text, sentence after Eq. (1)"},{"comment":"The captions of Fig. 2 and Fig. D.1 report scaling exponents b = 0.8 ± 0.1 and b = 0.6 ± 0.1 (and 0.8 ± 0.1, 0.5 ± 0.1 in Fig. D.1) without stating the fit range or confidence intervals; please specify the fitting procedure and the NA range used, and clarify whether the errors are statistical or systematic.","section":"Fig. 2 caption"},{"comment":"The claim that FGM is robust to 'global dephasing of arbitrary strength λ' is stated with only the order O(D^{-1}) error; since Appendix E shows that local depolarization affects Fmax at first order, the distinction between global and local noise in this statement should be made explicit.","section":"Appendix B, Eq. (B1)"},{"comment":"The paper cites relevant work on randomized measurements and direct fidelity estimation, but does not cite the companion experimental paper that introduced the data in Ref. [28] beyond the data source; consider acknowledging that the data were not collected for the present protocol and that the measurement sequences were post-hoc re-used, which is relevant to the interpretation of Fig. 4.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful protocol paper with a clean derivation, but the advertised 'cross-platform' demonstration is absent, and the Appendix E robustness claim contains an actual counterexample (orthogonal states with overlapping marginals yield positive estimated fidelity under unitary mismatch). These are fixable in revision: the authors can either (a) provide an experimental or numerical two-device test with independent unitaries, or (b) substantially weaken the claims and re-frame the paper as a protocol proposal with single-platform validation. The central identity and resource analysis are solid and should be preserved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core of this paper is real. Eq. (2) — estimating Tr[ρ1ρ2] between two platforms from cross-correlations of randomized measurements — is genuinely new relative to the same group's purity protocol [17], and the Appendix A proof via Weingarten calculus is clean and standard. I checked the derivation; it holds. The scaling analysis is useful, too: b ≈ 0.6–0.8 for product and Haar-random states versus b ≥ 2 for full tomography. The resource comparison is qualitative, since the exponents are numerical fits with ±0.1 error, but the direction is right.\n\nThe soft spot that matters is in the robustness claims. The main text says unitary errors and depolarization 'decrease the estimated fidelity and do not lead to false positives.' Appendix E establishes that only for ρ1 = ρ2. Their own Eq. (E6) contains a marginal-overlap term that produces exactly the failure mode a verification protocol cannot afford. Take ρ1 = |00⟩⟨00| and ρ2 the Bell state (|01⟩+|10⟩)/√2. Exact overlap is zero, but each single-qubit marginal has overlap 1/2, so the estimated overlap is positive at order η² and Fmax comes out ≈ 2η² (or η² with one-sided error). That is a false positive, and Fig. E.1(a) cannot see it because it only tests identical states. The stress-test note is correct on this; the fix is to restrict the claim to the equal-state case or to characterize the protocol as conservative only there.\n\nSecond, the experiment is not cross-platform. The experiment-theory comparison on re-used 10-qubit trapped-ion data from [28] is a legitimate proof of principle, and the paper honestly calls the experiment-experiment part 'same experiment.' But splitting one dataset into E1 and E2 yields a self-consistency check of the estimator, not evidence that two devices can be compared. The title and framing overreach what was demonstrated.\n\nMinor: the t=0 fidelity loss in Fig. 3 shows unitary implementation errors are significant in this platform; the authors are open about it, which is to their credit, but it reinforces how much the error analysis matters.\n\nWho this is for: anyone doing NISQ verification, randomized measurements, or fidelity estimation. Eq. (2) and the scaling analysis are worth having in the literature. The paper deserves serious refereeing — the core is publishable, but the Appendix E claim needs to be fixed and the experimental section reframed before it goes out.","headline":"Solid core: Eq. (2) is new and correctly derived, but the Appendix E 'no false positives' claim is false for states with overlapping marginals, and the experiment is one platform, not two.","tokens_in":96943,"tokens_out":5517,"would_cite":true,"duration_ms":56173,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A fidelity between two unknown quantum states on separate devices can be measured from shared random measurements and classical communication alone.","keywords":["cross-platform verification","quantum fidelity estimation","randomized measurements","unitary 2-designs","density matrix overlap","mixed state fidelity","trapped ion quantum simulator","intermediate scale quantum devices"],"falsifier":"Prepare the same known product state, such as $|0\\cdots 0\\rangle$, on both platforms and run the protocol. If the inferred $F_{\\max}$ is substantially below unity and drops further when each random unitary is replaced by two consecutive random unitaries, then systematic unitary errors and decoherence during the measurement basis are contaminating the estimate, and the protocol is measuring the unitary mismatch as much as the states.","tokens_in":95925,"feed_emoji":"⚛️","tokens_out":10729,"duration_ms":102055,"temperature":0.7,"pith_summary":"This paper establishes a practical way to compare two quantum states prepared on different experimental platforms without assuming what the states are and without sending quantum information between the platforms. It shows that if both devices apply the same classically communicated random local unitaries and then measure in the computational basis, the overlap $\\mathrm{Tr}[\\rho_1\\rho_2]$ of the two states is encoded in the cross-correlations of their outcome statistics; the same data also gives the purities $\\mathrm{Tr}[\\rho_i^2]$, so a mixed-state fidelity $F_{\\max} = \\mathrm{Tr}[\\rho_1\\rho_2]/\\max\\{\\mathrm{Tr}[\\rho_1^2], \\mathrm{Tr}[\\rho_2^2]\\}$ can be reported. This gives a verification tool for intermediate-scale devices that is exponentially cheaper than full quantum state tomography. The authors demonstrate the method on 10-qubit entangled states produced in a trapped-ion simulator, comparing experiment with theory and one experimental run with another.","feed_headline":"Cross-check two quantum devices with shared random measurements","feed_subtitle":"No quantum link needed: shared random unitaries reveal the fidelity of 10-qubit states.","key_machinery":"The workhorse is the cross-correlation estimator in Eq. (2): the weighted sum $\\sum_{s_A,s'_A}(-d)^{-D[s_A,s'_A]} P^{(1)}_{U_A}(s_A) P^{(2)}_{U_A}(s'_A)$, averaged over shared local unitary 2-designs. A unitary 2-design is an ensemble whose first two moments match uniformly random unitaries, which lets the average be evaluated analytically. Under the ensemble average the two-copy operator built from this sum collapses to the swap operator, leaving exactly $\\mathrm{Tr}[\\rho_1\\rho_2]$; the same machinery yields purities as autocorrelations. The protocol thus converts a quantum overlap into ordinary statistical correlations of measurement outcomes, which can be compared between two platforms using only classical communication.","core_discovery":"The central claim is that the cross-platform fidelity of two unknown, possibly mixed states can be inferred directly from randomized measurements. For subsystems $A_1,A_2$ with equal size $N_A$, let $U_A = \\otimes_{k=1}^{N_A} U_k$ be a product of independent local random unitaries drawn from a unitary 2-design, and let $P^{(i)}_{U_A}(s_A)$ be the probability that device $i$ observes outcome string $s_A$ after applying $U_A$. The paper proves $$\\mathrm{Tr}[\\rho_{1,A_1}\\rho_{2,A_2}] = $d^{{N_A}}$ \\sum_{s_A,s'_A} (-d)^{-D[s_A,s'_A]} \\overline{ $P^{{(1)}}$_{U_A}(s_A) $P^{{(2)}}$_{U_A}(s'_A) },$$ where $D[s_A,s'_A]$ is the number of positions at which the two outcome strings differ and the overline is the ensemble average over the random unitaries. Setting the two indices equal recovers the purities from the same formula, and normalizing the overlap by the larger purity defines $F_{\\max}$. The proof uses a two-design averaging identity to turn the averaged two-copy operator into the swap operation, so that the cross-correlation sum becomes the trace product of the two density matrices.","pith_inferences":["One consequence the authors leave implicit is that the shared random-unitary set could become a common reference for a network of different quantum devices, certifying two machines as equivalent without either being trusted as the ground truth, provided both can implement the same declared unitary set.","The sensitivity to unitary mismatch is itself a diagnostic resource: measuring the same known state with one versus two concatenated random unitaries separates state fidelity from unitary calibration error, because the estimated fidelity drops when the unitary path lengthens.","The scaling analysis suggests that an adaptive measurement-allocation scheme, guided by bootstrap error estimates, could reduce the total budget below the quoted $2^{bN}$ bound when partial prior knowledge of the states is available.","The same estimator could be used to track a system's memory of its own earlier state, since the fidelity of a state with its own version at a later time decays slowly in the disordered setting the paper studies."],"forward_implications":["Two quantum devices in different places and times can be checked against each other with only classical communication of random unitaries and outcomes, with no quantum link required.","A quantum simulator can be verified against a classical simulation of its target state; the paper reports experiment-theory fidelities above 0.6 even after 5 ms of many-body dynamics on 10 qubits.","The measurement budget scales as roughly $2^{bN_A}$ with $b\\approx 0.6$-$0.8$ for $N_A$-qubit subsystems, well below the $2^{2N_A}$ scale of full state tomography, making tens of qubits accessible.","The same data yields any fidelity that depends only on overlap and purities, including the geometric-mean fidelity $F_{\\mathrm{GM}}$, which is first-order insensitive to local depolarizing noise.","By splitting one experimental dataset into two 'experiments', the method also benchmarks the reproducibility of a single device; the authors find experiment-experiment fidelities higher than experiment-theory fidelities."],"supporting_citations":[{"why":"provides the single-system randomized-measurement purity estimator that Eq. (2) generalizes to cross-platform overlaps.","marker":"[17]"},{"why":"defines the unitary 2-designs from which the shared local random unitaries are sampled.","marker":"[29]"},{"why":"supplies the formal treatment and constructions of unitary $k$-designs that the protocol's sampling assumption rests on.","marker":"[30]"},{"why":"contains the Appendix A proof of Eq. (2) via the two-design averaging identity.","marker":"[31]"},{"why":"provides the trapped-ion randomized-measurement data used for the 10-qubit proof-of-principle fidelities.","marker":"[28]"},{"why":"sets the full quantum state tomography scaling baseline that the protocol's smaller measurement budget is compared against.","marker":"[19]"},{"why":"supplies the bootstrap resampling technique used to assign statistical error bars to the estimated fidelities.","marker":"[33]"}],"fun_headline_variants":["Randomized measurements verify two quantum devices at once","No quantum link: verify 10-qubit state across platforms","Cross-platform fidelity from random local measurements","Measure overlap of two quantum states with random bases","Shared random unitaries reveal cross-device fidelity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that both devices can implement the same declared random unitaries accurately enough that the measured cross-correlations are dominated by the difference between the two states rather than by the difference between the two implementations of the random unitaries.","fun_headline_variants_meta":{"raw":{"variants":["Randomized measurements verify two quantum devices at once","No quantum link: verify 10-qubit state across platforms","Cross-platform fidelity from random local measurements","Measure overlap of two quantum states with random bases","Shared random unitaries reveal cross-device fidelity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00059,"raw_usage":{"total_tokens":2761,"prompt_tokens":931,"completion_tokens":1830,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":547,"completion_tokens_details":{"reasoning_tokens":1758}},"tokens_in":547,"tokens_out":1830,"duration_ms":12020,"temperature":1.0,"reasoning_tokens":1758,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:22:45.489360+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Prepare the same known product state, such as $|0\\cdots 0\\rangle$, on both platforms and run the protocol. If the inferred $F_{\\max}$ is substantially below unity and drops further when each random unitary is replaced by two consecutive random unitaries, then systematic unitary errors and decoherence during the measurement basis are contaminating the estimate, and the protocol is measuring the unitary mismatch as much as the states.","supporting_citations":[{"cited_title":"Gross, K","cited_arxiv_id":null,"evidence_quote":"defines the unitary 2-designs from which the shared local random unitaries are sampled."},{"cited_title":"Gross, Y","cited_arxiv_id":null,"evidence_quote":"sets the full quantum state tomography scaling baseline that the protocol's smaller measurement budget is compared against."},{"cited_title":"Efron and G","cited_arxiv_id":null,"evidence_quote":"supplies the bootstrap resampling technique used to assign statistical error bars to the estimated fidelities."}],"review_version":1}