{"id":"81cca739-59ea-47f8-85f6-a847aae22877","arxiv_id":"2606.28833","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"Neyman shot allocation driven by three closed-form GP pair sensitivities plus a 50% uniform floor yields 10–21% RMSE gains over uniform allocation for quantum-kernel Gaussian processes in the moderate-budget regime.","lead":"The paper shows how to spend limited quantum-circuit shots unevenly across kernel matrix entries so Gaussian-process regression stays accurate. The method cuts test error 10–21% versus uniform sampling on moderate budgets and works for several GP tasks when the kernel still has usable pair-to-pair differences.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged floor calibration.","rationale":"The central claim is an empirical performance statement under a concrete allocator (predictive-coupling sensitivity + ρ=0.5 floor + calibrated jitter). The supporting theory (Propositions 1, 2, 4–6) correctly identifies Neyman optimality for the estimation-error objective and the necessity of coverage; the only non-derived design choice that the claim depends on is the numerical value of the floor. The reader already flags this precisely, and the paper itself documents the calibration. No stronger load-bearing flaw (e.g., an incorrect gradient, a missing term that changes the ranking of pairs, or an unacknowledged concentration regime that would erase the reported quantum-kernel gains) appears after a full-manuscript reading. Therefore the CONDITIONAL verdict with high confidence remains appropriate; the concrete floor-sweep test would simply tighten or loosen that conditionality without altering the overall assessment.","tokens_in":24502,"tokens_out":606,"duration_ms":5633,"concrete_test":"Re-run the four UCI benchmarks (n_tr=200, 10 seeds, B=10^6) with ρ∈{0.3,0.4,0.5,0.6} while keeping the pre-committed |α_i α_j| sensitivity; if the paired gain of gp_alpha vs uniform remains ≥10% and p<0.05 for every ρ≥0.4, the floor choice is robust; if the gain collapses or loses significance for any ρ≤0.4, the headline claim is more fragile than stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest_assumption already isolates the softest point: Proposition 3 only proves that some positive uniform floor is necessary to drive |S| to 0 (so that the operator-norm lower bound on missing-entry perturbation vanishes); it does not derive that ρ=0.5 is theoretically adequate for the condition numbers and budgets used. The paper is explicit that ρ=0.5 is an empirical default chosen from the dense-synthetic ablation (Figure 6) and that the jump from the coverage breakpoint ≈0.1 to 0.5 is a shot-quality margin. Because the headline 10–21% RMSE claim is measured under exactly this fixed floor, and because the same floor is held constant across UCI, quantum-kernel, and extension experiments, the claim is conditional on that calibration transferring. No deeper internal inconsistency or hidden assumption that would invalidate the Neyman derivation or the three sensitivities was found; the mathematics is elementary and the negative concentration result is reported honestly.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper extends active quantum kernel acquisition (AQKA) from classification to Gaussian process regression under a finite shot budget. Each Gram entry is a Bernoulli estimate; the authors derive three closed-form pair-level sensitivities (predictive coupling |α_i α_j|, leave-one-out residual, and marginal-likelihood gradient), plug them into a Neyman minimum-variance allocation rule, and add a high uniform coverage floor (ρ=0.5) plus inference-time jitter to avoid catastrophic over-concentration when the warm-up kernel is noisy. On four UCI regression benchmarks and two synthetic RBF+Bernoulli studies (n_tr=200), the allocator yields 10–21% test-RMSE gains over uniform allocation in the moderate-budget regime; gains transfer to genuine ZZ/Pauli-Z kernels on quantum-natural data (−13–15% at low budget, paired p<0.05) and to Bayesian quadrature, heteroscedastic GP, hyperparameter learning, and multi-output Cokriging. On UCI features embedded into a ZZ map the gain vanishes, consistent with exponential concentration. Six short propositions justify the Neyman rule, posterior error propagation, spectral heterogeneity, the coverage floor, and the status of each sensitivity relative to the exact estimation-error objective.","tokens_in":24764,"tokens_out":1599,"duration_ms":23919,"significance":"If the empirical gains hold under the stated conditions, the work is a useful and timely extension of shot-budgeted quantum kernel methods from 0/1 classification to full GP regression, where inverse-amplified quantities (predictive variance, log-det, NLL) make shot allocation more consequential. Strengths include: (i) elementary but correctly stated closed-form sensitivities with explicit status relative to exact Neyman weights (Propositions 4–6, Corollary 1); (ii) an honest negative result on UCI+ZZ embeddings and an explicit catastrophic low-budget regime with floor ablation; (iii) transfer experiments to sparse VFE, BO, streaming, and multi-output settings; (iv) paired t-tests and NLL/Frobenius diagnostics that separate “better where it matters” from overall kernel accuracy. The free parameters (ρ, jitter, warm-up fraction) are documented and held fixed after ablation. The contribution is incremental relative to prior AQKA classification work but non-trivial for GPs, and the theory-to-experiment mapping is unusually careful for the area.","major_comments":[{"comment":"Proposition 3 only establishes a necessary coverage condition: some positive uniform floor is required so that the missing-pair set S vanishes and the operator-norm lower bound on Δ does not stay positive. The operational choice ρ=0.5 (and the jump from the coverage breakpoint ≈0.1 to 0.5) is calibrated solely by the dense-synthetic ablation in Figure 6 / Result 3 and then frozen across UCI, quantum-kernel, and extension experiments. Because the headline 10–21% RMSE claim is measured under exactly this fixed floor, the central empirical claim is conditional on that calibration transferring. The manuscript should either (a) provide a broader ablation over σ_n, condition number, and n/B ratios that practitioners will encounter, or (b) state more sharply that ρ is a free hyperparameter requiring validation on a small pilot budget, and report sensitivity of the UCI gains to ρ∈{0.3,0.5,0.7}.","section":null},{"comment":"Result 10 and the abstract: the positive transfer claim (−13–15% on ZZ/Pauli-Z, p<0.05) is restricted to planted-sparse quantum-natural data at small q, while Study 1 (UCI features through ZZ) is essentially null and trends worse at larger q. Both findings are scientifically valuable, but the abstract and introduction currently lead with the positive transfer and relegate the null result to a consistency remark with Thanasilp et al. For a methods paper aimed at near-term quantum kernels, the scope condition—“gain exists only when pair-level sensitivity heterogeneity survives concentration and depolarizing noise”—should be stated as a primary takeaway, not a caveat, so that readers do not over-generalize the UCI RBF+Bernoulli gains to arbitrary quantum feature maps.","section":null}],"minor_comments":[{"comment":"Method vs Proposition 5: the factor-of-2 discrepancy between g_marg in Eq. (14) and the single-coordinate form in Proposition 5 is explained in the text, but a single consistent convention (and a one-line note that the Neyman weights absorb the constant) should appear at first use of sens_marg to avoid reader confusion.","section":null},{"comment":"Jitter (Eq. 19 and Theory paragraph): the manuscript correctly notes that the code uses √n·variance rather than the Wigner-motivated √(n·variance) and that both are clipped near 0.5 in the operating range. Consider moving this honesty into the main Algorithm 1 caption or a short remark so that implementers do not treat j as theoretically fixed.","section":null},{"comment":"Table 1: several headline gains are non-significant (e.g., energy at 2e5, california at 1e6 p=0.086). The abstract’s “10–21%” range is accurate for the moderate-budget sweet spot but should be qualified as “up to / on 3 of 4 datasets at B=1e6” to match the table.","section":null},{"comment":"Figure 1 caption and Result 1: at very low budget all AQKA variants underperform uniform; this is important and already analyzed, but a single sentence in the abstract (“gains appear only above ~50 shots/pair”) would set expectations correctly for hardware users.","section":null},{"comment":"Notation: A := K+σ_n^{2}I is standard, but b_ti and eta (sparse) appear without a compact symbol table; a short notation paragraph or table would help.","section":null},{"comment":"Related work: the concurrent Miroszewski (2026) and Xu et al. (2026) AQKA classification papers are cited; a one-sentence contrast on why GP needs a higher floor than classification (already in the intro) could be repeated briefly in Related Work for readers who skip the introduction.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is carefully written and unusually honest about negative results and free parameters. The main risk for the journal is over-claiming transfer to “genuine quantum kernels” when the positive quantum results are on synthetic quantum-natural data and the real-feature ZZ embedding fails. If the authors tighten the abstract/scope language and expand the floor-sensitivity discussion, this is a solid methods contribution suitable for the venue. No integrity or citation-pattern concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a practical, well-executed extension of the recent AQKA classification work to Gaussian-process regression. The new pieces are three closed-form pair sensitivities (predictive coupling |αiαj|, LOO residual, marginal-likelihood gradient), the observation that inversion-based predictors need a much higher uniform coverage floor than classification, and the transfer experiments to genuine ZZ/Pauli-Z kernels plus four downstream GP tasks.\n\nWhat they do well is keep the math elementary and honest. Propositions 1–6 are standard resolvent / Neyman arguments; they correctly flag that |αiαj| is only the rank-1 specialization of the exact predictive-MSE weight under M=yy⊤, that the LOO form drops a bounded correction, and that Proposition 3 only proves some positive floor is necessary, not that ρ=0.5 is theoretically adequate. The experiments match: paired t-tests on four UCI sets, floor ablation, NLL-vs-Frobenius diagnostics, sparse VFE, N-scaling, and an explicit null result when UCI features are pushed through a concentrating ZZ map. Gains of 10–21 % RMSE in the moderate-budget regime (~50–250 shots/pair) and −13–15 % on quantum-natural data are reported with p-values and seeds; the free parameters (ρ, jitter, warm-up) are held fixed after the ablation rather than re-tuned per figure.\n\nThe soft spot is exactly the one the stress-test flags: ρ=0.5 is an empirical default chosen on dense synthetic data. The coverage lower bound only forces ρ ≳ n^{2}/(2B) ≈ 0.1; the jump to 0.5 is a shot-quality margin. Because every headline number uses that fixed floor, the claim is conditional on the calibration transferring. That is a real but proportionate caveat, not a load-bearing flaw. No code is shipped, which is annoying for a methods paper but not fatal given the fully specified protocol.\n\nThis is for people who already care about shot-budgeted quantum kernels or about making GPs work under expensive kernel evaluations. It will not change asymptotic sample complexity, but it is a usable engineering improvement with clean theory and careful negatives. I would send it to referees; the contribution is real and the writing is careful enough that a serious review will improve it rather than sink it.","headline":"Clean, usable extension of Neyman shot allocation from classification kernels to GP regression, with honest negatives and solid empirics; the only real soft spot is the empirically fixed 50% floor.","tokens_in":25337,"tokens_out":631,"would_cite":true,"duration_ms":6357,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Non-uniform shot allocation for quantum kernels cuts Gaussian-process test error by 10–21% in the moderate-budget regime.","keywords":["quantum kernels","Gaussian process regression","shot allocation","Neyman allocation","kernel estimation","active learning","quantum machine learning"],"falsifier":"Run the same UCI and synthetic protocols with the uniform floor set to zero or 10% at moderate budgets (~50 shots per pair); if the allocator still matches or beats uniform RMSE without the high floor, the necessity claim fails.","tokens_in":25417,"feed_emoji":"⚛️","tokens_out":945,"duration_ms":8538,"temperature":0.7,"pith_summary":"Quantum kernels on near-term hardware must estimate every Gram-matrix entry from a finite number of circuit shots. Prior work showed that, for classification, spending more shots on kernel entries that most affect the loss improves accuracy under a fixed budget. This paper extends the same idea to Gaussian-process regression, where the quantities that matter—predictive means and variances, the log-determinant, and the marginal likelihood—are more tightly coupled to kernel error than a simple sign. The authors derive three closed-form pair sensitivities (predictive coupling, leave-one-out residual, and marginal-likelihood gradient), feed them into a Neyman minimum-variance allocation rule, and protect the scheme with a high uniform-coverage floor so that a noisy warm-up estimate cannot starve important pairs. On four UCI benchmarks and controlled synthetic studies the allocator yields 10–21% lower test RMSE than uniform allocation across the moderate-budget window; the same gains appear for genuine ZZ and Pauli-Z kernels on quantum-natural data and transfer to Bayesian quadrature, heteroscedastic regression, hyperparameter learning and multi-output Cokriging. When UCI features are forced into a concentrating ZZ kernel the advantage vanishes, showing that the method only helps when the kernel still carries usable pair-level heterogeneity.","feed_headline":"Smart shot allocation cuts quantum-GP error 10–21%","feed_subtitle":"Weighting circuit shots by GP sensitivity beats uniform allocation under the same total budget","key_machinery":"The three pair-level sensitivities—predictive coupling |α_i α_j|, leave-one-out residual, and marginal-likelihood gradient—plugged into the Neyman rule s*_ij ∝ |g_ij| √[K_ij(1-K_ij)] together with a high uniform-coverage floor justified by a Frobenius lower bound on missing-entry perturbation.","core_discovery":"A Neyman-style allocation that weights each quantum-kernel entry by one of three closed-form GP sensitivities, protected by a 50% uniform floor, reduces test RMSE by 10–21% relative to uniform shot allocation in the moderate-budget regime, and the improvement transfers to genuine quantum kernels on data that retain pair-level heterogeneity as well as to several standard GP downstream tasks.","pith_inferences":["A fully online multi-round re-estimation loop that updates sensitivities after each small batch of shots could push the useful regime down by another order of magnitude in total budget.","Composing the entry-wise allocator with classical leverage-score row selection would further cut the number of distinct circuits that must be submitted.","The same Neyman weights could be used as an acquisition score for choosing which new data points to label when both labels and shots are scarce."],"forward_implications":["Practitioners can obtain lower test error from the same total circuit shots when fitting GPs with quantum kernels in the 50–250-shots-per-pair window.","The same sensitivities and floor can be ported, with only rank-1 adjustments, to sparse inducing-point GPs and deep GPs.","Shot budgets for Bayesian optimization, Bayesian quadrature and multi-output Cokriging can be reduced by the same mechanism whenever the underlying kernel retains pair-level heterogeneity.","When a quantum feature map enters the exponential-concentration regime, non-uniform allocation ceases to help, giving a practical diagnostic for whether a kernel is still usable."],"fun_headline_variants":["GP-sensitivity shot weights cut quantum-GP RMSE 10-21%","Neyman allocation trims quantum kernel GP error 10-21%","Pair sensitivities guide shots for 10-21% GP RMSE gain","Task-aware shots beat uniform on quantum GP by 10-21%","Uniform-floor Neyman rule lifts quantum-GP accuracy 10-21%"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That a fixed 50% uniform floor, chosen by ablation on synthetic data, is enough to keep a noisy warm-up sensitivity estimate from catastrophically over-concentrating shots for the budgets and condition numbers that arise in practice.","fun_headline_variants_meta":{"raw":{"variants":["GP-sensitivity shot weights cut quantum-GP RMSE 10-21%","Neyman allocation trims quantum kernel GP error 10-21%","Pair sensitivities guide shots for 10-21% GP RMSE gain","Task-aware shots beat uniform on quantum GP by 10-21%","Uniform-floor Neyman rule lifts quantum-GP accuracy 10-21%"]},"model":"grok-4.5","effort":"low","cost_usd":0.003928,"raw_usage":{"total_tokens":1296,"prompt_tokens":866,"num_sources_used":0,"completion_tokens":83,"cost_in_usd_ticks":39280000,"prompt_tokens_details":{"text_tokens":866,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":347,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":866,"tokens_out":83,"duration_ms":3319,"temperature":1.0,"reasoning_tokens":347,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T11:14:25.982183+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same UCI and synthetic protocols with the uniform floor set to zero or 10% at moderate budgets (~50 shots per pair); if the allocator still matches or beats uniform RMSE without the high floor, the necessity claim fails.","supporting_citations":[],"review_version":3}