{"id":"0d53fbbb-555b-4567-aae1-7748cd101233","arxiv_id":"2505.22514","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"GPU-based simulated bifurcation closes the reported quantum annealing scaling advantage on Sidon-28 QUBO instances, with robust classical scaling on larger problems.","lead":"A classical solver running on GPUs matches or beats a D-Wave quantum annealer on the benchmark problems that were recently used to claim a quantum scaling advantage. The result suggests the reported speedup came from the choice of classical comparison method and from testing only small problems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Large-instance scaling may rest on uncertified ground states and unpublished instances; the asymptotic claim is not yet supported.","rationale":"The reader's weakest assumption is exactly the risk on the large-instance leg. I agree. The paper's main abstract conclusion about QPUs relies more on the large-N scaling than on the small-N comparison alone; the small-N comparison already shows parity, but 'unlikely to demonstrate supremacy' is an extrapolation that needs a valid large-N baseline. The absence of instance release and certification details makes the finding non-reproducible. This does not overturn the small-N parity result, which appears solid given the public instances and consistent runtime accounting, so the verdict remains CONDITIONAL rather than REJECT. The proposed test would settle the matter by checking whether the scaling exponents are robust under exact E0 certification.","tokens_in":72,"tokens_out":5406,"duration_ms":82693,"concrete_test":"Ask the authors to release the L=20..80 instances and Gurobi logs (or run Gurobi with a 0% MIP gap and a much larger time limit on L=20, 40, 80, or a subset of 10 instances). Independently recompute E0 and the probability p used in TT_epsilon, then re-fit α. If any E0 is not certified optimal, or α shifts by more than the reported ±error, the asymptotic scaling claim is unsupported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The abstract and main text conclude that current QPUs are unlikely to show a scaling advantage in approximate optimization, based on two legs: the direct comparison on N≤1322 instances and a new scaling analysis on N=2380–38320 'Sidon-28' instances (Fig. S2, α≈1.5–1.7). The second leg is load-bearing for the asymptotic claim, but it assumes two unverified conditions. First, the large instances are generated identically to the original Ref. [1] instances; the SM states 'we consider the same Sidon-28 instances' yet immediately says 'For each L, we consider 10 random instances', indicating new draws. No seed, generation script, or the instances themselves are provided, so the distribution cannot be checked. Second, Eq. (1) defines TT_epsilon relative to E0 'certified by Gurobi solver', but the Methods give no optimality gap, time limit, or certificates, and the SM is silent for L=20–80. Obtaining a rigorous optimality proof for random sparse QUBOs of N≈38k is non-trivial; if the reported E0 is only an incumbent (upper bound), then epsilon|E0| is too small relative to the true gap, p is overestimated, and TT_epsilon is underestimated, artificially lowering α. This would directly invalidate the 'robust classical performance' conclusion. The small-instance closing-the-gap result (Fig. 1) is less affected, since those instances are public and E0 was certified in the prior work, but the paper's broad final claim is not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reassesses the recent quantum-annealing scaling-advantage claim of Munoz-Bauza and Lidar by benchmarking a GPU implementation of the Simulated Bifurcation Machine (SBM) on the same Sidon-28 QUBO instances (N=142 to N=1322). Using the time-to-epsilon metric of Ref. [1], it reports SBM scaling exponents comparable to or lower than D-Wave QAC and PT-ICM, and it argues that the small-instance exponents are unstable with respect to runtime accounting and problem size. The paper then presents new large-instance results (N=2380 to N=38320) with scaling exponents alpha approximately 1.5 to 1.7 and concludes that current-generation quantum annealers are unlikely to demonstrate true supremacy in approximate optimization under operationally meaningful conditions.","tokens_in":10549,"tokens_out":4526,"duration_ms":55887,"significance":"If correct, the paper would substantially qualify a published PRL advantage claim and would strengthen the case that classical GPU-based chaotic solvers provide a competitive benchmark for approximate QUBO optimization. The small-instance benchmark has clear strengths: it uses instances from the original public dataset, 100 independent runs per instance, bootstrap error bars, and a transparent distinction between total runtime and pure GPU runtime. The large-instance extrapolation, however, is the load-bearing component of the asymptotic conclusion, and its current presentation does not fully support it. The manuscript also makes a useful methodological point that scaling exponents obtained from small sizes depend strongly on runtime accounting and on the choice of epsilon, which deserves to be preserved in any revised version.","major_comments":[{"comment":"The asymptotic claim alpha approximately 1.5-1.7 rests on TT_epsilon values computed with Eq. (1) using E0 'certified by Gurobi solver', but the paper provides no optimality certificates, solver time limits, or optimality gaps for the L=20 to L=80 instances (N up to 38320), and it does not release these instances or a generation script. If E0 is only an upper bound, then p(E <= E0 + epsilon|E0|) is overestimated and TT_epsilon is underestimated, which can artificially lower the fitted alpha and invalidate the robustness conclusion. In addition, the SM says 'we consider the same Sidon-28 instances' but immediately describes '10 random instances' per size, so the relationship to the original Ref. [1] instances is unclear. The authors should either certify exact optimality for the large instances and make the instances and certificates available, or explicitly restrict the asymptotic claim to the N <= 1322 regime.","section":"Supplemental Material, 'Scaling in large instance regime'; Fig. S2; Eq. (1)"},{"comment":"The SBM hyperparameters N_s and N_r are optimized on the same instances used to produce the scaling exponents, with no cross-validation or hold-out set. Because TT_epsilon is minimized over (N_s, N_r) separately for each instance size and each epsilon value, the reported exponents reflect in-sample tuning, and the degree of overfitting is unknown. Before accepting the 'closing the gap' claim as evidence of a general classical scaling property, the authors should report a validation split, nested optimization, or some other guard against selection bias in the hyperparameters.","section":"Fig. 1, Fig. S2, and Methods"},{"comment":"The large-instance scaling is based on only 10 instances per size, yet the paper reports no per-size spread, confidence intervals, or instance-level TT_epsilon values for this regime, in contrast to the small-instance analysis where bootstrap errors are shown. With ten samples, the median and the fitted exponent are sensitive to a few outlier instances; the authors should report per-size distributions or bootstrap intervals for [TT_epsilon]_Med in Fig. S2 before describing the large-instance scaling as 'robust'.","section":"Fig. S2 and accompanying SM text"}],"minor_comments":[{"comment":"The phrase 'true supremacy' is stronger than the evidence presented; a narrower statement about scaling advantage on Sidon-28-type QUBO instances would better match the scope of the benchmark.","section":"Abstract and Discussion"},{"comment":"The coefficient 0.7 in Delta(t) = 0.7 t/T is introduced without justification or sensitivity analysis; a brief discussion of its role or a reference to the quantized-SBM literature would help.","section":"Eq. (3) and Methods"},{"comment":"There are typographical errors ('Advantege series', 'problem siszes' in Ref. [30]) and the statement that the large instances are 'the same Sidon-28 instances' should be reworded, since new random draws are described.","section":"Supplemental Material"},{"comment":"The manuscript does not state where the code, the large-instance generation script, or the instance files will be deposited; a data availability statement is needed, especially since the large-instance results are not derived from the public Harvard Dataverse dataset.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The paper is a direct response to a high-profile PRL advantage claim, and its negative conclusion will attract attention. The key missing piece is reproducibility of the large-instance leg: exactness certificates, instance generation details, and data availability. If the authors cannot supply these, the asymptotic conclusion should be softened accordingly. I do not see a circularity or misconduct concern, but the in-sample hyperparameter optimization should be addressed as a correctness-risk issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is the paper. It is an empirical benchmark that directly challenges Munoz-Bauza and Lidar's PRL claim of a quantum scaling advantage in approximate optimization. The authors run the simulated bifurcation machine (SBM) on the same public Sidon-28 QUBO instances used in the PRL, and show that GPU-based SBM matches or beats the D-Wave QAC annealer in time-to-epsilon across the four epsilon values, whether you count total runtime or pure GPU time. If it holds, that is a real result. The data presentation is candid: they report total and pure GPU runtimes, bootstrap error bars, and reproduce the prior numbers.\n\nWhat is actually new is not SBM itself, nor the ternary discretization from Zhang and Han, but the empirical closing of the gap on the exact benchmark instances, plus a first scaling study on much larger instances (N up to 38320) of the same family. That large-instance part is where my main concern sits. The paper claims the large instances are 'the same Sidon-28 instances' but then says they drew 10 random instances per size for L=20-80. No instances, seeds, or generation scripts are provided. More importantly, the time-to-epsilon definition requires the exact ground state energy E0, and the text only says it is 'certified by Gurobi solver' with no optimality gap or time limit. Certifying random sparse QUBOs at N=38320 is nontrivial. If the reported E0 is only an upper bound, epsilon is effectively larger than stated, success probabilities are overestimated, and time-to-epsilon is underestimated, artificially lowering the scaling exponent. That would weaken the 'robust classical performance' conclusion and the final claim about QPUs.\n\nThat said, this concern does not break the central small-instance result, which uses the public instances from the PRL with previously certified ground states. The hyperparameter tuning on the same instances used for scoring is a legitimate caveat, but it is the same practice as in the PRL and not disqualifying. The conclusion is broader than one problem family and one epsilon range can fully support, but as a benchmark result the paper is honest and clear. It deserves a serious referee, and I would bring it to the reading group.","headline":"A credible SBM benchmark closes the reported quantum annealing scaling gap on the original instances, but the large-instance extrapolation rests on ground-state certifications and unpublished instances that need scrutiny.","tokens_in":11141,"tokens_out":2741,"would_cite":true,"duration_ms":28448,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A GPU-based classical solver matches or beats the quantum annealer's time-to-epsilon scaling on Sidon-28 QUBO instances, so the reported quantum scaling advantage does not survive a different classical baseline.","keywords":["quantum annealing","QUBO","simulated bifurcation machine","time-to-epsilon","scaling advantage","Sidon-28 instances","GPU computing"],"falsifier":"Recompute the large-instance TT-$\\varepsilon$ curves with independently certified ground-state energies (for example, by an exact branch-and-bound solver or by matching lower bounds), and also recompute the small-instance comparison with a runtime definition that includes all quantum-device programming and readout times; if either change makes the quantum-annealing exponent smaller than SBM's, the paper's conclusion that current quantum annealers show no clear advantage would be overturned.","tokens_in":10034,"feed_emoji":"⚙️","tokens_out":12250,"duration_ms":119494,"temperature":0.7,"pith_summary":"This paper tests a recent claim that quantum annealing achieves a scaling advantage in approximate optimization of QUBO problems. By replacing the original classical baseline (parallel tempering) with a GPU-implemented simulated bifurcation machine, the authors find that the classical solver matches or outperforms the quantum annealer's time-to-epsilon on the same Sidon-28 instances. They argue that the small instance sizes used previously (up to about 1300 variables) cannot support asymptotic conclusions, because the measured scaling exponent depends strongly on how runtime is counted. Extending to instances with up to 38,320 variables, they report classical scaling exponents around 1.5-1.7 that are stable across the optimality gaps tested. If accepted, the conclusion is that current-generation quantum annealers are unlikely to show a clear, operationally meaningful scaling advantage on this class of problems; the paper points to sparse problem classes as the plausible terrain for future quantum advantage once hardware overheads shrink.","feed_headline":"GPU solver closes reported quantum annealing scaling gap","feed_subtitle":"On Sidon-28 QUBO instances, simulated bifurcation matches the quantum annealer's time-to-epsilon.","key_machinery":"The yardstick is the time-to-epsilon metric, $\\mathrm{TT}_\\varepsilon = t_f \\log(0.01)/\\log(1 - p_{E \\le E_0+\\varepsilon|E_0|})$ (Eq. 1), which converts a solver's measured running time and success probability into a single effective time to reach a given relative optimality gap $\\varepsilon$. The contender is the simulated bifurcation machine, a nonlinear Hamiltonian dynamical system whose equations (Eqs. 2-3) couple variables through Ising interactions, drive the system through a bifurcation point with a linear schedule $a(t)=t/T$, and impose an inelastic wall at $|q_i|=1$; a time-dependent threshold discretizes the interaction to the sign of the variable. Chaotic sensitivity lets many replicas run in parallel on GPUs, and the hyperparameters — the number of steps $N_s$ and number of replicas $N_r$ — are optimized per instance size. This TT-yardstick is used to compare SBM, PT-ICM, and two quantum-annealer sampling methods on identical instances.","core_discovery":"The paper's central claim is that the time-to-epsilon scaling exponent of the simulated bifurcation machine on the exact Sidon-28 instances used in Ref. [1] is comparable to or smaller than the exponent of quantum annealing with error correction (QAC) for optimality gaps 0.75% to 1.25%, even when all CPU-GPU communication overhead is included. It further claims that exponents fitted to the small instance set are not reliable indicators of asymptotic behavior; the same SBM data, fit over $N = 2380$ to $38320$, gives exponents between $1.5$ and $1.7$ that vary only weakly with $\\varepsilon$ and with runtime accounting. The conclusion the authors draw is that the previously reported quantum-classical gap is an artifact of the chosen classical baseline and of finite-size effects, and that current QPUs do not exhibit genuine supremacy in approximate discrete optimization under operationally meaningful conditions.","pith_inferences":["The large-instance result is only as strong as the ground-state certification: if the reported $E_0$ values for $N$ up to $38320$ are not actually exact, the fitted exponents could be systematically too low and the asymptotic conclusion would weaken.","The runtime-accounting sensitivity shown here suggests a simple test for future advantage claims: report total wall-clock time including all solver overheads, and compare against at least one chaotic-dynamics solver and one thermal solver.","Because SBM exponents improve with more GPUs and with newer classical hardware, the classical baseline will keep moving; quantum advantage claims should specify the classical hardware generation, as happened with random-circuit sampling.","The sparse-instance, large-$\\varepsilon$ window that the paper mentions as favorable to QPUs is also one where SBM needs no modification, so the prediction of future quantum advantage there is conditional on hardware-overhead reductions the paper does not quantify."],"forward_implications":["SBM with one GPU and total runtime matches or beats QAC for all tested optimality gaps on the small Sidon-28 instances, so the reported quantum advantage does not hold against a different classical baseline.","Scaling exponents from $N \\le 1322$ are sensitive to whether runtime is counted as wall-clock or pure GPU time, so small instances cannot anchor asymptotic claims about quantum advantage.","At larger sizes ($N = 2380$ to $38320$), SBM's time-to-epsilon scaling lies in $1.5 < \\alpha < 1.7$ and is nearly independent of $\\varepsilon$ in the tested range, providing a stable classical baseline for future quantum-device comparisons.","Current-generation quantum annealers are therefore unlikely to demonstrate a clear scaling advantage on QAC-type QUBO problems under the time-to-epsilon accounting used here.","Sparse problem classes are identified as the promising place to seek genuine future quantum advantage, conditional on reduced hardware overheads."],"supporting_citations":[{"why":"The quantum annealing scaling-advantage claim this paper reassesses; supplies the PT-ICM, U3, and QAC data and the small-instance benchmark.","marker":"[1]"},{"why":"Defines PT-ICM, the classical baseline the paper argues was an unrepresentative choice.","marker":"[20]"},{"why":"Introduces the simulated bifurcation machine, the solver whose time-to-epsilon scaling is measured.","marker":"[22]"},{"why":"Introduces the discretized-interaction version of SBM used in the paper's equations.","marker":"[23]"},{"why":"Supplies the quantum annealing correction (QAC) scheme whose scaling advantage is being questioned; the paper compares SBM against QAC data.","marker":"[26]"},{"why":"Supplies the Sidon-28 instances shared with Ref. [1], making the comparison instance-for-instance.","marker":"[27]"},{"why":"Certifies the exact ground-state energies $E_0$ used to define the optimality gap in Eq. (1).","marker":"[28]"}],"fun_headline_variants":["Classical solver matches quantum annealer scaling on QUBO","Simulated bifurcation closes quantum annealing scaling gap","Quantum scaling advantage refuted by classical dynamical solver","GPU solver matches annealer time-to-epsilon on QUBO","No quantum scaling edge for approximate optimization"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The large-instance scaling exponents that anchor the asymptotic conclusion assume that the Sidon-28 instances with $L=20$ to $L=80$ are drawn from the same distribution as the original instances and that the reported ground-state energies are exact up to $N=38320$.","fun_headline_variants_meta":{"raw":{"variants":["Classical solver matches quantum annealer scaling on QUBO","Simulated bifurcation closes quantum annealing scaling gap","Quantum scaling advantage refuted by classical dynamical solver","GPU solver matches annealer time-to-epsilon on QUBO","No quantum scaling edge for approximate optimization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000261,"raw_usage":{"total_tokens":1588,"prompt_tokens":935,"completion_tokens":653,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":577}},"tokens_in":551,"tokens_out":653,"duration_ms":6537,"temperature":1.0,"reasoning_tokens":577,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:05:42.004179+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the large-instance TT-$\\varepsilon$ curves with independently certified ground-state energies (for example, by an exact branch-and-bound solver or by matching lower bounds), and also recompute the small-instance comparison with a runtime definition that includes all quantum-device programming and readout times; if either change makes the quantum-annealing exponent smaller than SBM's, the paper's conclusion that current quantum annealers show no clear advantage would be overturned.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the simulated bifurcation machine, the solver whose time-to-epsilon scaling is measured."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the discretized-interaction version of SBM used in the paper's equations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the quantum annealing correction (QAC) scheme whose scaling advantage is being questioned; the paper compares SBM against QAC data."}],"review_version":1}