{"id":"af6a6966-07f2-4ed9-b6cb-d35475631404","arxiv_id":"2502.08060","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"QAEMCMC uses quantum annealing outputs as Metropolis-Hastings proposals and claims faster mixing for the N=10 Sherrington-Kirkpatrick model, but the reported advantage rests on oracle-tuning tau and one scaling table contradicts the text.","lead":"This paper proposes a hybrid sampling method that uses quantum annealing to generate candidate states inside a classical Markov-chain Monte Carlo loop, and reports faster convergence on a small spin-glass benchmark. The practical impact depends on whether the per-instance tuning used in the experiments can be replaced by a fixed schedule, which is not shown.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"QA advantage is shown only when the annealing time is tuned per instance against the exact spectral gap (Fig. 2); a fair fixed-schedule test, the one a practical sampler would face, is missing.","rationale":"The reader and I identify the same weakness. The abstract and introduction promise that QAEMCMC 'yields significant spectral gaps and faster convergence,' but the only protocol that produces the headline gaps optimizes τ against the exact spectral gap. This makes the comparison unfair in a specific, testable way: the classical updates have no analogous per-instance hyperparameter tuned to the spectral gap. A fixed-schedule benchmark would resolve whether the advantage is an artifact of the oracle. I also note the Table 2 inconsistency, which further undermines confidence in the numerical reporting, but it is secondary; the oracle dependence is the load-bearing issue. My recommendation is to leave the reader's REJECT verdict unchanged: the concern is substantive, and the paper currently lacks the fixed-schedule evidence needed to support ACCEPT or CONDITIONAL.","tokens_in":10845,"tokens_out":5515,"duration_ms":51899,"concrete_test":"Re-run the complete N=10/T=1 and N-scaling benchmarks on the same 100 instances with a small set of fixed annealing times (e.g., τ = 0.01, 1, 10, 100, 1000) chosen before any spectral-gap computation, and also with an adaptive schedule based only on observable energy estimates; report spectral gaps, mixing-time proxies, and TV exponents. If the QA update no longer beats uniform and local updates for these practical schedules, the central claim rests on the per-instance spectral-gap oracle.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central numerical evidence for faster mixing is generated under an oracle protocol. In Results, Fig. 2, the authors state: 'τ is tuned to maximize the absolute spectral gap by Optuna for T < 10.' That is, for each SK instance the proposal distribution is selected by maximizing the very quantity used to claim acceleration, so the reported gap is an envelope over annealing times rather than the performance of a definite algorithm. The system-size scaling in Fig. 3/Table 1 does not state how τ is chosen, so the reported exponent α=0.254 for QA may inherit the same per-instance optimization. The only non-oracle dynamical evidence (Figs. 4–5) uses one selected N=10 instance, and Table 2 is internally inconsistent with the text: the τ_MCS exponents duplicate Table 1's spectral-gap values and show QA with the smallest α, contradicting the claim that 'QA update gives larger exponents.' Because no code or data are released, the reader cannot check whether the advantage survives a fixed, a priori annealing schedule. The paper therefore does not support the central claim as a practical algorithm; it supports only the existence of a favorable τ per instance.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes QAEMCMC, in which a quantum annealer generates proposal states inside a Metropolis-Hastings chain, with the acceptance step enforcing detailed balance so that the target Gibbs-Boltzmann distribution is preserved. The method is benchmarked on small Sherrington-Kirkpatrick instances (N = 10 and a few larger sizes) against classical local and uniform updates. The authors report larger absolute spectral gaps, faster convergence of the energy estimator, and smaller total variation distance to the target distribution, and they extract scaling exponents for the spectral gap and TV distance as functions of system size.","tokens_in":11064,"tokens_out":5267,"duration_ms":40408,"significance":"If the reported advantage survives a non-oracle choice of the annealing schedule, QAEMCMC would be a conceptually interesting way to accelerate MCMC in frustrated spin systems, and it connects naturally to the existing quantum-enhanced MCMC literature. The paper correctly identifies the M-H correction as the mechanism that preserves the target distribution, and the use of QA's diabatic outputs as proposals is worth exploring. However, as presented, the evidence is not yet convincing: the central spectral-gap comparison optimizes the annealing time per instance against the very quantity being reported, the dynamical evidence rests on a single selected instance, and Table 2 contains an internal inconsistency. Because these issues directly affect the main claims, the current significance is limited.","major_comments":[{"comment":"The reported QA spectral gap is obtained by per-instance optimization of the annealing time: 'In the QA update, τ is tuned to maximize the absolute spectral gap by Optuna for T < 10.' This makes the comparison an oracle comparison: the QA proposal is selected after seeing the exact spectral gap, while the classical proposals are fixed. For large systems, the spectral gap cannot be computed, so the reported advantage (including the exponent α=0.254 in Table 1) does not correspond to a practical algorithm. To support the central claim, the authors should report results under a fixed, a priori annealing schedule, or under an adaptive choice of τ that uses only quantities available to a running sampler. Without such a test, the paper only demonstrates that a favorable τ exists for each instance.","section":"Results, Fig. 2"},{"comment":"The system-size scaling in Fig. 3 does not state how τ is chosen for each N. If the same Optuna-based maximization is used, the fitted exponent inherits the oracle-selection problem. The caption and text should specify the τ-selection protocol for the scaling data; if τ is optimized per instance, the reported α is an envelope over optimized schedules, not the exponent of a definite algorithm.","section":"Results, Fig. 3 and Table 1"},{"comment":"The first numerical column of Table 2 reports α = 0.939, 0.855, 0.254 for uniform, local, and QA updates. These values are identical to the spectral-gap exponents in Table 1. For the claimed scaling TV(P_ex, Q_emp) ~ τ_MCS^{-α}, the QA value is the smallest, which means the slowest decay, contradicting the statement in the text that 'The QA update gives larger exponents than the local updates.' This internal inconsistency must be resolved: either the fits are wrong, or the text/table labels are wrong. The TV-scaling claim cannot be evaluated until this is corrected.","section":"Results, Table 2"},{"comment":"The dynamical simulations are performed on a single, deliberately selected instance ('we select a challenging problem instance used in Fig. 2'). The 'twelve times faster' energy convergence and the TV distance behavior are thus anecdotal. The authors should average over many instances and report the distribution of convergence times, especially given the strong instance-to-instance variation in the optimized τ shown in Fig. 2(b)-(c). As written, the dynamical evidence is not representative.","section":"Results, Figs. 4 and 5"}],"minor_comments":[{"comment":"Equation (6) appears to have typographical errors: the mixing-time bounds '(1 − δ −1)ln2 ε' and '−δ −1 ln (εminσσσ µ(σσσ ))' are not correctly typeset, and the second inequality should be checked against the standard bound in Levin and Peres. Please clarify.","section":"Methods, Eq. (6)"},{"comment":"There is a typo 'QA-ehnhanced' in the section heading; also, the text refers to 'Metoplized independent sampling' where the intended term is 'Metropolized.'","section":"Methods, 'QA-enhanced MCMC'"},{"comment":"The notation for the target distribution is inconsistent: Eq. (8) uses μ(σ), while the results use P_ex for the exact distribution and Q_emp for the empirical distribution. Please define these symbols explicitly at first use.","section":"Results, Fig. 5"},{"comment":"The data availability statement mentions datasets, but not code. Given that the numerical claims rely on a specific Optuna protocol, the authors should release the code or at least specify the search range, number of trials, and seeding for the τ optimization.","section":"Data Availability"}],"recommendation":"major_revision","confidential_remarks":"The oracle-tuning issue is the main substantive concern: unless the authors can demonstrate the advantage under a fixed or adaptively estimated annealing schedule, the central claim should be substantially weakened. The internal inconsistency in Table 2 also needs immediate correction. I would encourage the editor to request a thorough revision rather than reject outright, because the underlying idea is reasonable and the identified problems appear addressable with additional numerical experiments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nI read Arai & Kadowaki's QAEMCMC paper. Bottom line: the idea is a plausible, incremental twist on QEMCMC—use forward quantum annealing to generate Metropolis-Hastings proposals for SK-model sampling—but the numerical evidence as presented does not establish a practical speedup. The key issue is the oracle tuning of annealing time.\n\nWhat's new: direct use of QA output as M-H proposal with Q_QA(σ)=|⟨σ|ψ⟩|², preserving detailed balance by construction. That's a genuine variation on Layden et al.'s QEMCMC (ref 21) and close to Christmann et al.'s reverse-annealing variant (ref 56), which they cite. The paper is transparent about the protocol and the small N. The sample analysis (Hamming distance vs energy difference) is a nice way to explain why QA proposals might help.\n\nThe soft spot is load-bearing. In Fig. 2, τ is optimized per instance via Optuna to maximize the absolute spectral gap—the very quantity used to claim acceleration. That makes the reported gap an envelope over schedules, not the performance of a definite algorithm. A practical sampler can't compute the spectral gap for large N. The system-size scaling in Fig. 3 likely inherits the same tuning, though it isn't stated. The only non-oracle dynamical evidence (Figs. 4–5) uses a single selected N=10 instance, which can be cherry-picked. On top of that, Table 2's τ_MCS exponents duplicate Table 1's spectral-gap exponents and show QA with the smallest α, contradicting the text's claim that 'QA update gives larger exponents.' That's an internal inconsistency that should have been caught. No code or data is released, which makes checking anything harder.\n\nTo be fair: the proposed method is well-defined, the M-H correction is sound, and the paper acknowledges the small-size limitation. The per-instance tuning is disclosed in the text, not hidden. But the central claim—faster mixing than classical updates—is not supported for a fixed or adaptive schedule, and the table error undermines confidence.\n\nWho this is for: anyone working on quantum-assisted MCMC or hybrid sampling will want to know this exists, but they should read the numbers skeptically. It deserves a serious referee, but with heavy revision required: redo the benchmark under a fixed schedule (or at least report both fixed and tuned), fix the tables, and upload code/data. I'd lean toward rejecting the current version but inviting a resubmission along those lines.\n\nBest.","headline":"The M-H construction is sound and the QA proposal idea is a real incremental twist, but per-instance tuning of annealing time against the exact spectral gap makes the reported speedup an envelope, not a practical algorithm, and Table 2 is internally inconsistent.","tokens_in":11590,"tokens_out":2174,"would_cite":false,"duration_ms":18411,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["05.10.Ln","03.67.Ac"],"model":"deepseek-v4-flash","headline":"A quantum annealing sweep can serve as a proposal generator inside Markov-chain Monte Carlo, and on the Sherrington-Kirkpatrick spin glass it accelerates convergence by more than twelvefold while preserving the exact target distribution.","keywords":["quantum annealing","Markov chain Monte Carlo","Metropolis-Hastings","Sherrington-Kirkpatrick model","spectral gap","mixing time","spin glass","Boltzmann sampling"],"falsifier":"Run QAEMCMC on Sherrington-Kirkpatrick instances at sizes around N=16 to 20 with a fixed or adaptively estimated annealing time (no spectral-gap oracle), and compare the convergence of the energy estimator and total variation distance against local and uniform updates; if the QA chain is not faster under such a schedule, the central claim fails. A second check: monitor the acceptance probability of QA proposals as N grows, since a decay toward zero would neutralize the gap advantage.","tokens_in":10621,"feed_emoji":"⚛️","tokens_out":10656,"duration_ms":91391,"temperature":0.7,"pith_summary":"This paper asks whether a quantum annealer can be used not as a final sampler, but as a proposal generator inside a classical Markov-chain Monte Carlo loop. The authors insert the state produced by a short quantum annealing run into a Metropolis-Hastings acceptance test, which guarantees that the chain still converges to the exact Gibbs-Boltzmann distribution. On the Sherrington-Kirkpatrick spin-glass model at small sizes, they find that this QA-enhanced proposal gives a Markov chain with a much larger spectral gap than local or uniform spin updates, corresponding to dramatically faster mixing. The central quantitative claim is that the QA-update chain reaches a fixed accuracy in the energy estimator in about 618 Monte Carlo steps, more than twelve times faster than the best classical update tested. If this holds, quantum hardware could be used to accelerate sampling in spin-glass-like problems without changing the target distribution.","feed_headline":"Quantum annealing makes spin-glass sampling 12 times faster","feed_subtitle":"A hybrid sampler preserves the exact target distribution while quantum proposals jump between local minima.","key_machinery":"The central mechanism is the QA proposal distribution Q_QA(σ)=|⟨σ|ψ(τ)⟩|^2, obtained by evolving the uniform superposition under the transverse-field Hamiltonian H(s)=s H0 - (1-s) Σ σ^x for a time τ and measuring in the computational basis. The paper feeds this into the Metropolis-Hastings acceptance probability A(σ'|σ)=min(1, μ(σ')/μ(σ) · Q(σ|σ')/Q(σ'|σ)), so the detailed-balance condition is preserved and the stationary distribution is the target Boltzmann distribution. The absolute spectral gap δ of the resulting transition matrix carries the argument: the mixing time is bounded by the inverse gap, so a larger gap directly implies faster convergence. The paper computes δ as a function of temperature, system size, and annealing time τ, with τ tuned per instance to maximize δ.","core_discovery":"The discovery is that taking the output of a quantum annealing sweep, specifically the probability distribution over spin configurations at the end of a Schrodinger evolution under a transverse field that is slowly turned off, as the proposal distribution in a Metropolis-Hastings update yields a Markov chain whose spectral gap can be much larger than that of local or uniform updates on the SK model. Because the chain uses the standard acceptance probability, the stationary distribution remains exactly the Gibbs-Boltzmann target; the QA solver only biases which moves are proposed. The authors report that at temperature T=1, the spectral gap of the QA-update chain decays with system size N as roughly $2^{{-0.254 N}}$, versus $2^{{-0.939 N}}$ for uniform updates and $2^{{-0.855 N}}$ for local updates. On a hard instance, the QA chain reaches an energy error of 0.1 in about 618 Monte Carlo steps, while the local update needs about 7,920 and the uniform update about 21,232, an improvement of more than twelvefold.","pith_inferences":["The paper tunes the annealing time per instance by maximizing the exact spectral gap; whether the speed-up survives with a fixed or self-adapting schedule is the key open question, and it can be tested by re-running the dynamics with a single schedule chosen from the reported optimized-τ histograms.","Because the QA proposal here is independent of the current state, QAEMCMC is effectively a Metropolized independent sampler; this suggests the essential ingredient is a proposal distribution that is both close to the target and broad, which might also be approximated classically.","The paper's own caveat about unequal weights for degenerate ground states means the low-temperature fairness of the empirical distribution is not guaranteed by the large spectral gap alone and should be checked in practice.","A natural extension is to condition the QA proposal on the current state via reverse annealing; the paper notes this alternative but does not test it."],"forward_implications":["Any MCMC code that uses Metropolis-Hastings can incorporate a QA proposal without changing its target distribution, so the speed-up is portable to other Boltzmann samplers.","On the SK model at T=1, reaching an energy error of 0.1 takes about 618 Monte Carlo steps with the QA update versus 7,920 with the local update and 21,232 with the uniform update, a speed-up of more than twelvefold.","The spectral-gap decay exponent at T=1 is 0.254 for QA versus 0.939 for uniform and 0.855 for local updates, so the gap advantage widens with system size within the small sizes simulated.","QA proposals combine large Hamming distance from the current state with small energy difference, which explains the high acceptance probability and the ability to move between local minima.","The empirical distribution from the QA chain approaches the target Gibbs distribution faster in total variation distance, as shown by the earlier convergence visible in the dynamics simulations."],"supporting_citations":[{"why":"Supplies the quantum-enhanced MCMC template, generating proposals with quantum evolution and correcting with Metropolis-Hastings, which this paper adapts to quantum annealing.","marker":"[21]"},{"why":"Defines the Sherrington-Kirkpatrick model used as the benchmark for all spectral-gap and dynamics simulations.","marker":"[49]"},{"why":"Gives the acceptance criterion that guarantees the QA-update chain's stationary distribution matches the target.","marker":"[5]"},{"why":"Provides the mixing-time bound that connects the absolute spectral gap to convergence rate, the main performance metric.","marker":"[4]"},{"why":"Used to tune the annealing time per instance by maximizing the spectral gap, the procedure whose success underlies the reported advantage.","marker":"[53]"},{"why":"Establishes the Metropolized-independent-sampling form of the acceptance probability that applies when the QA proposal is independent of the current state.","marker":"[50]"}],"fun_headline_variants":["Quantum annealing boosts MCMC sampling 12x on spin glasses","Quantum proposal moves make MCMC converge faster","Hybrid sampler uses QA proposals to speed up MCMC","QA-enhanced MCMC widens spectral gaps, speeds sampling","Quantum-assisted MCMC achieves 12x speedup on SK model"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The advantage depends on choosing the annealing time individually for each instance by maximizing the exactly computed spectral gap, a procedure that cannot be carried out for the large systems where the speed-up would matter.","fun_headline_variants_meta":{"raw":{"variants":["Quantum annealing boosts MCMC sampling 12x on spin glasses","Quantum proposal moves make MCMC converge faster","Hybrid sampler uses QA proposals to speed up MCMC","QA-enhanced MCMC widens spectral gaps, speeds sampling","Quantum-assisted MCMC achieves 12x speedup on SK model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000717,"raw_usage":{"total_tokens":3172,"prompt_tokens":843,"completion_tokens":2329,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":459,"completion_tokens_details":{"reasoning_tokens":2248}},"tokens_in":459,"tokens_out":2329,"duration_ms":30736,"temperature":1.0,"reasoning_tokens":2248,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T10:58:07.580283+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run QAEMCMC on Sherrington-Kirkpatrick instances at sizes around N=16 to 20 with a fixed or adaptively estimated annealing time (no spectral-gap oracle), and compare the convergence of the energy estimator and total variation distance against local and uniform updates; if the QA chain is not faster under such a schedule, the central claim fails. A second check: monitor the acceptance probability of QA proposals as N grows, since a decay toward zero would neutralize the gap advantage.","supporting_citations":[{"cited_title":"& Kirkpatrick, S","cited_arxiv_id":null,"evidence_quote":"Defines the Sherrington-Kirkpatrick model used as the benchmark for all spectral-gap and dynamics simulations."},{"cited_title":"& Peres, Y .Markov Chains and Mixing Times","cited_arxiv_id":null,"evidence_quote":"Provides the mixing-time bound that connects the absolute spectral gap to convergence rate, the main performance metric."}],"review_version":1}