{"id":"d219f78f-0d4d-4b34-a056-e8d3ab85c004","arxiv_id":"2411.12131","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A commercial SDK claims to simulate 53-qubit Sycamore circuits on 32GB RAM with an average XEB of 0.678, but the benchmark is weakly supported and partly self-referential.","lead":"This paper reports that a classical simulator, Quantum Rings SDK, can sample from Google's 53-qubit Sycamore random circuits on a 32GB machine, achieving an average linear cross-entropy benchmarking (XEB) score of 0.678. The authors claim this is the highest fidelity observed for these circuits, but the simulation method is not described and the score for the largest circuits appears to be measured against the simulator's own output.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline XEB of 0.678 is partly self-referential: for n=51 and n=53 the paper uses the SDK's own amplitudes as the ideal distribution because Google's amplitude files are unavailable, so the fidelity measure is not independently grounded.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing concern: the unavailability of independent ideal amplitudes for the largest circuits makes the reported XEB a self-comparison. This concern is central because the headline average 0.678 includes these circuits and because the claim of 'highest fidelity' depends on comparing XEB values computed against a fixed external ground truth. The proposed test is concrete and can be performed with existing data in the repository, without requiring the unavailable Google amplitudes for n=51 and n=53. Other issues, such as the undisclosed simulation method and the inappropriate comparison to noisy hardware, are serious but secondary; the self-referential XEB alone is sufficient to invalidate the central claim as stated. Therefore the preprint should be rejected as a scientific claim, unless the authors provide an independent reference for the large-circuit XEB values.","tokens_in":7946,"tokens_out":4152,"duration_ms":42604,"concrete_test":"For a circuit where Google's amplitudes are available (e.g., n=50, m=14, pattern EFGH), compute F_XEB twice: once using Google's amplitudes as p(x) and once using the SDK's own amplitudes as p(x), from the same sample set. If the SDK-based F_XEB substantially exceeds the Google-based F_XEB, the self-referential methodology inflates the score; applying the same comparison to n=51 and n=53 via an independent tensor-network amplitude computation (e.g., Pan & Zhang 2021) would then confirm the inflation.","verdict_should_be":"REJECT","load_bearing_attack":"Equation (2) defines F_XEB relative to an ideal distribution p(x). For n=51 and n=53, Figure 4 marks the range where Google's amplitudes are unavailable, forcing the paper to compute p(x) from the Quantum Rings SDK's own output. Yet the same SDK generated the sampled bitstrings. If the sampler and the reference distribution are both derived from the same approximate amplitudes, the resulting XEB measures self-consistency, not fidelity to the ideal quantum state. An approximate simulator can inflate its XEB by using its own biased distribution as the 'ideal' reference. The reported value for n=53 is 0.622, which is far below the value 1 expected if the SDK's amplitudes were exact and the samples were drawn from them; this inconsistency suggests that the sampling and reference distributions are not properly matched, or that the approximation is not controlled. Because the average 0.678 includes these self-referential points, it is not a valid comparison to Google's hardware XEB values, and the conclusion that this is the 'highest fidelity observed to date' is unsupported. This is the most load-bearing flaw: removing the n=51 and n=53 circuits, or recomputing them against an independent ground truth, could materially change the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports classical simulation of Google's random-circuit-sampling circuits from the Dryad dataset using the Quantum Rings SDK on machines with only 32 GB of memory. The authors sample 500,000 bitstrings for most circuits and 2,500,000 bitstrings for n=51 and n=53, compute linear cross-entropy benchmarking (XEB) using Eq. (2), and report an average F_XEB of 0.678, which they interpret as strong correlation with ideal quantum simulation and as the highest fidelity observed to date. They also report time-to-first-sample for the largest n=53 circuits.","tokens_in":8147,"tokens_out":14590,"duration_ms":146440,"significance":"If valid, the central claim would be a notable practical demonstration that large random-circuit-sampling circuits can be classically simulated on modest hardware, with useful implications for quantum circuit development and debugging. The paper has genuine strengths: it uses Google's published QASM circuits, makes source code and amplitude files publicly available on GitHub, checks the sample distribution against the Porter-Thomas prediction, and for circuits up to n=50 uses Google's own amplitude files as an external XEB reference. However, the headline average is compromised by the treatment of the largest circuits, and the comparison with prior work is not quantified.","major_comments":[{"comment":"For n=51 and n=53 the reported XEB values are not independently grounded. Equation (2) defines F_XEB relative to an ideal distribution p(x_i), but Figure 4 explicitly marks n=51/53 as the range for which Google's amplitudes are unavailable, while the text reports F_XEB=0.622 at n=53 and includes these circuits in the average 0.678. The only available source of \"ideal\" amplitudes in this range is the Quantum Rings SDK itself, so both the sampled bitstrings and the reference distribution are outputs of the same approximate simulator; this measures self-consistency, not fidelity to the ideal quantum state. The fact that the reported n=53 value is 0.622 rather than approximately 1 shows that the reference and sampling distributions are not even properly matched, further confirming that the metric is not interpretable. The manuscript nowhere states what reference was used for n=51 and n=53, nor does it provide an independent validation for those circuits.","section":"Section III, Figure 4, Eq. (2)"},{"comment":"The text and the figure caption contradict each other about whether Google's amplitude files exist for n=51 and n=53. Section III says the 2,500,000-sample count for these circuits was chosen \"to ensure that the number of samples matched those in Google's amplitude files,\" which implies such files exist, while the Figure 4 caption says the dashed segment is the qubits range for which Google's amplitudes are unavailable. This distinction is load-bearing: if the files exist, the XEB should have been computed with that external reference and the figure should show the corresponding blue points; if they do not, the sentence about matching the sample count is unexplained and the circularity concern applies.","section":"Section III vs. Figure 4"},{"comment":"The claim that the average XEB of 0.678 \"exceeds the XEB values currently reported for the same circuits today\" and \"represents the highest fidelity observed to date\" is unsupported. No table, equation, or citation provides the prior XEB values for the same circuits from references [4-7,9,10,12,15]. Moreover, XEB for a classical simulator is not directly comparable to hardware XEB: a noiseless simulator would yield F_XEB=1 by definition, whereas hardware XEB is suppressed by gate errors. The authors should either provide a quantitative comparison with prior classical simulation results for these exact circuits or remove the superlative claim.","section":"Section V (Conclusion) and Abstract"},{"comment":"The manuscript gives no description of the simulation algorithm or its approximation error. It says only that the SDK \"output[s] the amplitudes of each measurement\" and that circuits were executed in a Python environment. Exact statevector simulation of 53 qubits would require roughly 64 PB of memory, far beyond the 32 GB used, so the simulator must be approximate. The reported average F_XEB of 0.678, which is well below the F_XEB=1 expected from an exact simulator, further indicates a substantial deviation from the ideal distribution. Without knowing whether the method is a tensor-network, low-rank, or heuristic approximation, and without error bounds, the XEB result cannot be interpreted as evidence of \"effective simulation,\" and the experiment cannot be independently reproduced from the text.","section":"Section IV"}],"minor_comments":[{"comment":"The Dryad dataset is dated \"June 13, 2022\" in Section III and \"June 23, 2022\" in Section IV for the same reference [7]; please reconcile the dates.","section":"Sections III and IV"},{"comment":"The name \"Artu et al.\" should be \"Arute et al.\"; additionally, the citation given as [5] is Boixo et al., so the attribution of the error-gate XEB argument should be checked.","section":"Section II"},{"comment":"The captions of Figures 2 and 3 do not identify which particular circuit (n, m, pattern, seed) produced the plotted distribution, so the Porter-Thomas validation is not reproducible.","section":"Figures 2 and 3"},{"comment":"Figure 5 reports time-to-first-sample for p='ABCDCDAB' and m=12-20, while the XEB study uses p='EFGH' and m=14; the relationship between the timing experiment and the XEB experiment should be stated explicitly.","section":"Figure 5"},{"comment":"The phrase \"the SDK was used to output the amplitudes of each measurement\" is ambiguous about whether the simulator produces the full 2^n amplitude vector or only probabilities for the sampled bitstrings; please clarify.","section":"Section IV"},{"comment":"In Eq. (2), p(x_i) is never explicitly defined as the ideal probability |<x_i|psi_U>|^2, which is particularly important here because the choice of reference distribution is the central methodological issue of the paper.","section":"Eq. (2)"}],"recommendation":"major_revision","confidential_remarks":"The first and third authors are affiliated with Quantum Rings Inc., the vendor of the SDK, and the benchmarking code resides in the company's public repository; the SDK itself is proprietary and requires a license key. This in itself is not disqualifying, but it strengthens the need for an independent, explicitly described amplitude reference for the largest circuits. The self-referential XEB for n=51 and n=53 is a fundamental flaw in the current evidence, and the superlative claims in the abstract and conclusion should not appear until the XEB is recomputed against a genuine external reference or those data points are removed and the claims adjusted accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know that this paper is a product benchmark for the Quantum Rings SDK, not a research advance. The simulation method is proprietary and never described, so the scientific content is limited to the XEB numbers. What it does well: it uses Google's public Sycamore circuits, provides code and amplitude data on GitHub, and for n<=50 it computes XEB against Google's amplitude files as an external reference, which is the right thing to do. The Porter-Thomas plots look consistent.\n\nThe soft spot is the n=51 and n=53 circuits. Google's amplitudes are unavailable there, so the paper uses the SDK's own amplitudes as the ideal reference. That makes the XEB self-referential: it measures how consistent the SDK is with itself, not how close it is to the ideal quantum state. The reported n=53 XEB of 0.622 is far below the 1.0 you'd expect if the SDK's amplitudes were exact and sampling were faithful, so either the reference amplitudes are approximate or the sampling is mismatched. Either way, the average 0.678 loses its meaning once those points are included.\n\nThe paper also overclaims. It compares its XEB to Google's hardware XEB, but a classical simulator should be compared to exact classical simulation, which gives XEB very close to 1. And \"highest fidelity observed to date\" is nonsense if you don't compare to prior classical simulators like Pan-Zhang or Huang et al.\n\nI agree with the reader's rejection. Without a disclosed method and an independent reference for the largest circuits, the headline result isn't supported. That said, this could be sent to a serious referee, because the issues are important and the paper might be salvageable if the authors recompute the XEB for n=51,53 against a known-good reference and describe what the simulator actually does. I wouldn't cite it, but I might bring it to a reading group as a case study in how not to benchmark a simulator.","headline":"Benchmark of a proprietary simulator; headline XEB is not independently grounded for the largest circuits.","tokens_in":8712,"tokens_out":4631,"would_cite":false,"duration_ms":48719,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims the Quantum Rings SDK can classically simulate 53-qubit Sycamore random circuit sampling circuits on 32 GB hardware, reaching an average linear XEB score of 0.678 that the authors call the highest observed for these…","keywords":["Sycamore circuits","random circuit sampling","linear cross-entropy benchmarking","classical simulation","quantum simulator","Porter-Thomas distribution","53-qubit circuits","XEB fidelity"],"falsifier":"Take the $n=51$ and $n=53$ circuits, compute reference amplitudes with an independent exact or high-precision approximate method such as tensor-network contraction, then recompute the linear XEB of the SDK samples against that reference; if the resulting scores fall well below 0.678, the paper's central fidelity claim fails, while if they remain high, the self-comparison concern is resolved.","tokens_in":7678,"feed_emoji":"⚛️","tokens_out":6764,"duration_ms":62889,"temperature":0.7,"pith_summary":"The paper argues that a classical simulator running in the Quantum Rings SDK can handle the full 53-qubit, 14-cycle random circuit sampling circuits from the Sycamore experiment on machines with only 32 GB of RAM, while still producing samples that look close to ideal. The evidence is an average linear cross-entropy benchmarking score of 0.678 across circuit sizes, which the authors say exceeds the values reported earlier for the same circuits. If true, this would mean ordinary developer hardware can validate and debug large quantum circuits now, rather than waiting for fault-tolerant quantum computers. The paper also reports that the sampled probabilities closely follow the Porter-Thomas distribution, consistent with the expected quantum dynamics.","feed_headline":"A 32 GB RAM simulator scores 0.678 XEB on Sycamore circuits","feed_subtitle":"The result suggests developers can debug large quantum circuits without a supercomputer or quantum hardware.","key_machinery":"The load-bearing object is the linear cross-entropy benchmark, defined in the paper as $\\mathcal{F}_{\\mathrm{XEB}} = N \\left(\\frac{1}{k} \\sum_{i=1}^{k} p(x_i)\\right) - 1$, which measures how strongly the sampled bitstring probabilities correlate with the ideal circuit probabilities; it approaches 1 for noiseless sampling and 0 for uniform sampling. The authors use this score as a proxy for fidelity, following the standard relation $\\mathcal{F}_{\\mathrm{XEB}} \\approx (1-\\epsilon)^{\\#\\text{gates}}$. The simulator converts the published QASM circuit descriptions into amplitudes, samples bitstrings, and the samples are checked against the Porter-Thomas distribution as a secondary validation.","core_discovery":"On the paper's own terms, the discovery is that a universal quantum simulator can reproduce the output distribution of the 53-qubit Sycamore random circuit sampling circuits with high fidelity under memory constraints typical of a developer machine. The authors compute the linear XEB score of equation (2) for circuits of $n=12$ through $n=53$ qubits, using the original experiment's published amplitude files as the ideal reference where available and the SDK's own amplitudes for $n=51$ and $n=53$. They report an average $\\mathcal{F}_{\\mathrm{XEB}} = 0.678$, a value of $0.622$ for the largest circuit at 2.5 million samples, and execution times they describe as reasonable. They conclude that this is the highest fidelity observed to date for these circuits.","pith_inferences":["The reported average depends partly on self-referential XEB for the 51- and 53-qubit circuits, where the simulator's own amplitudes serve as the ideal reference; an independent amplitude reference could move the headline 0.678, so the 'highest fidelity' claim should be read with that caveat.","The paper's own citation of work on limitations of linear cross-entropy suggests a testable extension: deliberately inject depolarizing noise into the simulator and check whether $\\mathcal{F}_{\\mathrm{XEB}}$ decays as predicted by the error model.","The same protocol could be pushed to the 60-qubit, 24-cycle random circuit sampling circuits from a later experiment; the scalability limits of the simulator are not tested by the 53-qubit dataset alone.","A direct comparison against an independent tensor-network-based reference computation for the largest circuits would convert the self-comparison into an external benchmark."],"forward_implications":["Simulating 53-qubit random circuit sampling circuits on 32 GB hardware means large-scale circuit debugging no longer requires a supercomputer or specialized hardware access.","Cross-entropy benchmarking can serve as a routine validation metric for classical simulators, not just for physical quantum processors.","The reported average score provides a new classical baseline for the 14-cycle Sycamore circuits that future simulation claims would need to beat.","If the execution times hold, developers can test commercial quantum algorithms at near-ideal fidelity before fault-tolerant quantum hardware exists."],"supporting_citations":[{"why":"Supplies the Sycamore random circuit sampling experiment and the quantum supremacy claim that this study targets as a benchmark.","marker":"[4]"},{"why":"Introduces random circuit sampling and the cross-entropy benchmarking method used to evaluate the simulator.","marker":"[5]"},{"why":"Provides the published QASM circuit files and amplitude files that form the experimental input and reference data.","marker":"[7]"},{"why":"Establishes the Porter-Thomas distribution check used to validate the sampled probabilities.","marker":"[8]"},{"why":"A prior classical simulation of Sycamore circuits whose XEB results are the baseline the paper claims to exceed.","marker":"[9]"},{"why":"Another classical simulation study of the same circuits, cited as a comparison point for reported fidelity values.","marker":"[10]"},{"why":"Reports a later 60-qubit random circuit sampling experiment, used in the comparison for observed fidelity values.","marker":"[15]"},{"why":"Cited for known limitations of linear cross-entropy as a fidelity measure, relevant to interpreting the reported scores.","marker":"[16]"}],"fun_headline_variants":["Classical simulator hits 0.678 XEB on Sycamore circuits","32GB RAM simulator reproduces Sycamore output at XEB 0.678","Sycamore RCS circuits simulated with 0.678 XEB on a dev machine","Quantum supremacy circuits run classically at 0.678 XEB","Desktop-class simulator achieves 0.678 XEB on Google Sycamore"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The unstated load-bearing premise is that for the 51- and 53-qubit circuits, where the published ideal amplitude files are missing, the simulator's own amplitudes can stand in as the 'ideal' reference for the XEB calculation; if that premise is false, the average score and the 'highest fidelity' claim weaken because the average includes those largest circuits.","fun_headline_variants_meta":{"raw":{"variants":["Classical simulator hits 0.678 XEB on Sycamore circuits","32GB RAM simulator reproduces Sycamore output at XEB 0.678","Sycamore RCS circuits simulated with 0.678 XEB on a dev machine","Quantum supremacy circuits run classically at 0.678 XEB","Desktop-class simulator achieves 0.678 XEB on Google Sycamore"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000739,"raw_usage":{"total_tokens":3256,"prompt_tokens":857,"completion_tokens":2399,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":473,"completion_tokens_details":{"reasoning_tokens":2294}},"tokens_in":473,"tokens_out":2399,"duration_ms":17662,"temperature":1.0,"reasoning_tokens":2294,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:52:47.825223+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the $n=51$ and $n=53$ circuits, compute reference amplitudes with an independent exact or high-precision approximate method such as tensor-network contraction, then recompute the linear XEB of the SDK samples against that reference; if the resulting scores fall well below 0.678, the paper's central fidelity claim fails, while if they remain high, the self-comparison concern is resolved.","supporting_citations":[{"cited_title":"A Blueprint for Demonstrating Quantum Supremacy with Superconducting Qubits","cited_arxiv_id":null,"evidence_quote":"Establishes the Porter-Thomas distribution check used to validate the sampled probabilities."}],"review_version":1}