{"id":"e9dac5f3-50f0-4e70-9d02-053589b0f988","arxiv_id":"2504.12575","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A benchmarking framework that maps quantum computer error rates against arbitrary circuit features and uses Gaussian process regression to build predictive performance models from sparse data.","lead":"Featuremetric benchmarking is a new way to measure quantum computer performance by tracking how error rates change with circuit features like size, depth, and two-qubit gate density. It generalizes the established volumetric benchmarking method and uses Gaussian process machine learning to predict performance from fewer test circuits.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that GP regression predicts circuit performance is validated only for in-distribution interpolation; the paper's own Section VI concedes out-of-distribution prediction is unlikely, so the central 'faithful models' claim is not established as stated.","rationale":"The paper is a solid methods contribution: it formalizes featuremetric benchmarking, provides a rigorous GP-based analysis pipeline, and demonstrates it on real hardware with plausible results. The central claim is well-supported for the narrow task of interpolating mean performance on the feature grid used in training; Section V and Fig. 7 convincingly show that a randomly selected 20-50% of feature values can reproduce the remaining mean values within roughly 3% absolute error. However, the load-bearing assumption for the broader claim of 'richer and more faithful models of quantum computer performance' is that the chosen features (width, depth, two-qubit density) are approximately predictive of performance for arbitrary circuits, not just for held-out points on the same grid. The evaluation protocol never leaves the convex hull of the training feature vectors, so the GP is tested only as an interpolator. Section VI explicitly concedes that the model is unlikely to predict circuits outside the training distribution. This means the abstract's phrasing overstates the result: the paper demonstrates efficient interpolation of a benchmark surface, not general predictive capability modeling. The residual within-feature-vector variance (Fig. 1b) further shows that individual circuit performance is not captured by the model, only the mean over the benchmark distribution. I agree with the reader that this warrants a CONDITIONAL rather than full ACCEPT. The concrete test, evaluating the GP on an excluded contiguous region of feature space, would determine whether the model has any extrapolative value. If it fails, the paper's contribution should be framed strictly as an interpolation and sampling-efficiency result, not as a predictive model for arbitrary circuits.","tokens_in":31995,"tokens_out":14625,"duration_ms":153042,"concrete_test":"Train the GP on the ibmq montreal data with all feature vectors of depth d > 64 excluded, and compute the mean absolute error on the excluded d > 64 region. If that error is more than twice the interpolation error reported in Fig. 7, the 'much less data' claim is an interpolation artifact and the model is not predictive for arbitrary circuits.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section V B and Fig. 7 evaluate the GP by training on random subsets of the same feature grid and testing on the remaining grid points; all test points lie inside the convex hull of the training feature vectors. This validates interpolation, not prediction of previously unexplored circuits. The caption of Fig. 6 explicitly says the model is 'very accurate at interpolation as opposed to extrapolation,' and Section VI concedes that the feature set is unlikely rich enough to predict circuits not drawn from the same distribution as the training data. The abstract's claim of 'richer and more faithful models of quantum computer performance' therefore rests on an extrapolation that is not demonstrated. In addition, the model is trained on the mean of K circuits per feature vector (Section V A), so it predicts the mean of the benchmark distribution, not s(c) for a specific circuit; the residual within-feature-vector variance (visible in Fig. 1b for Montreal) is not modeled. The load-bearing premise that a small feature set is approximately predictive is only partially supported: Fig. 5 shows the three-feature Forte1 data is 'not statistically consistent with monotonicity,' although the reported 4-sigma deviation is the minimum over many comparisons, so its statistical significance is overstated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Featuremetric benchmarking is proposed as a framework that generalizes volumetric benchmarking by characterizing a quantum computer's circuit-level performance as a function of a small set of circuit features (e.g., width, depth, two-qubit gate density), rather than only width and depth. The framework is formalized through definitions of circuit features, pseudo-features, feature vectors, and circuit-sampling distributions, and it is demonstrated in three cloud-executed experiments: a two-feature randomized-mirror-circuit benchmark on ibmq montreal, a three-feature randomized-mirror-circuit benchmark on ibmq algiers, and a three-feature random-Clifford process-fidelity benchmark on IonQ Forte1 with quasirandom (Sobol) feature-vector selection. For data analysis, the authors use standard and monotonic Gaussian process (GP) regression to learn the mean capability s̄(v) as a function of the feature vector, validating on held-out grid points from random train/test splits. They report mean absolute errors of roughly 3-5% when training on about half the data, demonstrate that a GP trained on 20-50% of the montreal data approximately reproduces the full volumetric benchmarking plot, and show that adding two-qubit gate density as a feature substantially improves predictive accuracy.","tokens_in":32269,"tokens_out":17718,"duration_ms":164806,"significance":"The framework is a natural, clearly formalized generalization of a widely used benchmarking methodology, and the demonstrated dependence on two-qubit gate density (Fig. 3 shows 51% vs. 7% mean success at (w,d)=(14,64) as ξ2Q goes from 0 to 1/4) gives concrete evidence that going beyond width and depth is valuable. The GP analysis is methodologically appropriate for low-dimensional interpolation, the held-out validation is sound, and the data-efficiency result for the two-feature volumetric case (Figs. 6 and 7) is a useful and honest quantitative contribution. The three real experimental datasets on IBM Q and IonQ systems of up to 27 qubits are a genuine strength, as is the authors' explicit Section VI statement of the extrapolation limitation. The paper's main weaknesses are the gap between the abstract-level 'faithful models' claim and the interpolation-only validation, the selected-extreme statistics used for the Fig. 5 monotonicity conclusion, and the lack of released data and code. With the claims scoped to interpolation, this will be a solid contribution to the quantum-benchmarking literature.","major_comments":[{"comment":"The abstract's claim that featuremetric benchmarking 'enables richer and more faithful models of quantum computer performance' is not supported for circuits outside the training distribution. The validation in Section V B trains the GP on random subsets of the measurement grid and tests on the remaining grid points, which is in-distribution interpolation, and the caption of Fig. 6 explicitly limits the 20%-data model's accuracy to 'interpolation as opposed to extrapolation.' Section VI itself concedes that 'it is unlikely that our current feature sets are rich enough for our GP models to accurately predict the performance of circuits that are not drawn from the same distribution as the training data' (with the manuscript's two typos in this sentence corrected). Because a capability model is of interest precisely for predicting untested circuits, the abstract and Section I should be scoped to interpolation within the explored feature ranges, or the authors should add an explicit out-of-distribution test (e.g., training on interior Sobol points and testing on boundary or previously unobserved feature vectors) to substantiate the stronger claim.","section":"Abstract; Section I; Section V B; Fig. 6; Section VI"},{"comment":"The statement that the Forte1 data is 'not statistically consistent with monotonicity' rests on a single value δ_v = (−4 ± 1)% that is described as 'four σ inconsistent with being non-negative,' with the argument that among 248 events one would not expect a 4-sigma event by chance. This significance claim is computed on the most extreme of the 248 δ values, and each δ is itself a minimum over many pairwise comparisons, so the quoted z-score is a selected-extreme statistic and the informal multiple-events argument does not account for that selection or for correlations among feature vectors. A bootstrap or permutation analysis of the maximum deviation is needed to support the quantitative claim; the qualitative conclusion that the three-feature data is much closer to monotonic than any two-feature projection is well supported by the histograms and will likely survive such an analysis. In addition, the monotonic GP of Fig. 7 (with sharpness parameter ν = 10^{-6}, Appendix C) is applied to this same dataset despite the manuscript's own finding that it is not statistically consistent with monotonicity; the paper should justify this modeling choice and discuss whether the hard monotonicity constraint distorts predictions where the observed response is locally non-monotonic.","section":"Section IV C; Fig. 5; Fig. 7; Appendix C"},{"comment":"Neither the processed benchmark data nor the analysis code is made available. The central quantitative claims — the data-efficiency of GP-based reconstruction of volumetric plots (Figs. 6 and 8), the mean-absolute-error curves (Fig. 7), and the monotonic-versus-regular GP comparison — cannot be reproduced or independently verified from the manuscript alone, and the monotonic GP with expectation propagation (Appendix C) involves numerical choices not fully specified in the text. A data availability statement and a release of the datasets and code (or a clearly identified repository) should be provided.","section":"Sections IV-V; Appendix C"}],"minor_comments":[{"comment":"The symbol η² is used for the signal variance in Eq. (B2) and again for the observation-noise variance in Eq. (B3), and the noise variance later appears as σ² I in Eq. (B6); please use distinct symbols for the signal and noise variances throughout the appendix.","section":"Appendix B"},{"comment":"The axis label 'Monoticity Score' and the legend text 'Monoticity in Process Fidelity Decay' contain a typo ('Monoticity' → 'Monotonicity'), and the caption should include the caveat, stated in the main text, that the three-feature data is not statistically consistent with monotonicity.","section":"Fig. 5"},{"comment":"The GP observation model (Eq. B3) is homoscedastic, but the per-feature-vector means carry different bootstrap standard errors (e.g., Fig. 3); please state whether heteroscedastic noise was modeled or explain why it was neglected.","section":"Section V A"},{"comment":"In the sentence quoting success probabilities at (w,d) = (14,64), the notation 'χ2Q' should be 'ξ2Q' to match the convention used elsewhere in the paper.","section":"Section IV B"},{"comment":"Reference [26] is cited informally as 'Tomesh et al.'; please use the numbered citation style of the rest of the manuscript.","section":"Section III A"},{"comment":"In the concession paragraph, 'although it unlikely that our current feature sets are rich enough for our GP models to accurate predict' should read 'although it is unlikely that ... to accurately predict'; the typo matters here because this sentence carries the principal scope limitation of the paper.","section":"Section VI"},{"comment":"The sentence 'We train a GP on two randomly-selected subsets ... and then predicted the mean success probabilities' mixes tenses; 'predicted' should be 'predict.'","section":"Section V B"}],"recommendation":"major_revision","confidential_remarks":"The experimental work appears genuine and the manuscript is well written, but the gap between the abstract's strong claims and the demonstrated interpolation-only validation should be closed in revision; I would also ask the editor to consider whether the journal requires data/code availability for benchmarking papers, since none is currently provided. The paper builds directly on the authors' prior capability-learning and mirror-circuit work (Refs. [22, 50-52, 59-62]); the novelty relative to Refs. [52] and [59-60], which already develop predictive capability models, should be stated more explicitly, for instance by emphasizing the feature-based benchmarking framework and the volumetric data-efficiency result as the new contributions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful, practical methods paper. It does what it says: it generalizes volumetric benchmarking to arbitrary circuit features and shows that Gaussian process regression can reconstruct a volumetric capability plot from roughly 20% of the data. The framework is clearly formalized (features, pseudo-features, sampling distributions), and the experiments on three real systems are described honestly. The paper deserves a serious referee.\n\nWhat's actually new: the formalization is mostly a natural extension of volumetric benchmarking plus the prior capability-learning work, but the GP-based analysis for sparse featuremetric data is a genuinely useful addition. Monotonic GP is not new in ML, but applying it here and showing the trade-off against data size is real. The continuous volumetric plot from partial data is a nice practical outcome.\n\nSoft spots: the abstract says 'richer and more faithful models,' but the evidence supports accurate interpolation within the sampled feature grid, not prediction for circuits drawn from a different distribution. The paper's own Section VI concedes that the chosen feature set is unlikely to support out-of-distribution prediction. That is not a hidden flaw, but it should curb the abstract's language. Second, the monotonic GP is applied even though Figure 5 shows the three-feature data is not statistically consistent with monotonicity, and the 2D projections are strongly non-monotonic. The authors are honest about this, but it means the monotonicity prior is doing real work and the model is smoothing over non-monotonic residuals. Third, the 4-sigma statement in Figure 5 is the minimum over many comparisons, so its significance is overstated as written. Finally, no code or data shipped; for a benchmarking methods paper that is a concrete gap, because the GP training procedure, including the EP approximation, would be non-trivial to reproduce exactly.\n\nNone of these are load-bearing. The central framework holds up; the limitations are mostly about scope and presentation. This paper is for people designing or analyzing quantum benchmarks, not for the general quant-ph reader. I'd bring it to a reading group and cite it if I worked in benchmarking. Recommendation: send to peer review, with a request to soften the abstract, add code/data, and fix the 4-sigma interpretation.","headline":"A practical, honest methods paper that generalizes volumetric benchmarking to arbitrary circuit features and shows GPs can reconstruct capability plots from sparse data, though the abstract oversells out-of-distribution prediction.","tokens_in":32797,"tokens_out":1771,"would_cite":true,"duration_ms":20248,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Featuremetric benchmarking claims that quantum computer performance on circuits can be modeled from a small set of circuit features, and shows that Gaussian process regression can reconstruct volumetric benchmark plots from a fraction of…","keywords":["featuremetric benchmarking","volumetric benchmarking","circuit features","Gaussian process regression","capability learning","mirror circuits","two-qubit gate density","process fidelity"],"falsifier":"Run two structurally different circuit families that cover the same feature values on one device, train the Gaussian process on one family, and evaluate it on the other; if the mean absolute prediction error is no better than predicting the global average performance, the feature set is not predictive and the framework's central assumption fails.","tokens_in":31820,"feed_emoji":"⚛️","tokens_out":4946,"duration_ms":49951,"temperature":0.7,"pith_summary":"The paper introduces featuremetric benchmarking, a framework for measuring how well a quantum computer executes circuits as a function of chosen circuit features, such as width, depth, and two-qubit gate density, rather than only the two features used in volumetric benchmarking. It argues that this generalization yields richer and more faithful performance models, because no two features fully determine circuit error. The paper demonstrates the approach on cloud-access quantum computers of up to 27 qubits and shows that Gaussian process regression can reconstruct entire volumetric benchmarking plots from as little as 20 percent of the data. If the framework works as claimed, a modest batch of feature-parameterized circuits could produce predictive capability models for a quantum computer.","feed_headline":"Three circuit features map quantum computer performance","feed_subtitle":"A new benchmarking method reconstructs full performance maps from a fraction of the circuits.","key_machinery":"The central object is the feature vector $\\vec{f}=(f_1,\\dots,f_\\chi)$, a small set of computable circuit properties, together with the sampling distribution $P_{\\vec{v}}$ over circuits whose expected feature value equals $\\vec{v}$. Performance is summarized by a capability function $s(c)$, for example the success probability of mirror circuits or the process fidelity of random Clifford circuits. The paper's analytic engine is Gaussian process regression, both ordinary and monotonic, which interpolates the mean capability $\\bar{s}(\\vec{v})$ across feature space; the monotonic version enforces the expected decay of fidelity with increasing depth, width, and two-qubit density. Quasirandom Sobol sampling of feature vectors makes the experiment fill a $\\chi$-dimensional feature space efficiently.","core_discovery":"On its own terms, the paper's central claim is that the performance of a quantum computer on a circuit family is well approximated by a low-dimensional function of that circuit's features, so learning that function from measured circuits yields a predictive capability model. Concretely, it claims that with three features, namely width, depth, and two-qubit gate density, faithful performance models can be learned, and that Gaussian process regression can reproduce a volumetric benchmarking plot with high accuracy using as little as 20 percent of the data. The paper also establishes a formal framework: a featuremetric benchmark is defined by a feature vector, a rule for sampling circuits at each feature value, and a capability function such as success probability or process fidelity, with estimates obtained by running the sampled circuits.","pith_inferences":["A natural extension the authors leave implicit is active learning: because Gaussian processes supply uncertainty estimates, a benchmark could choose the next feature vector where model uncertainty is highest, reducing circuit counts further.","The framework invites learned features: instead of hand-picking width, depth, and density, one could use representation learning to find feature directions that minimize residual variance, making models more transferable across devices.","If the feature-predictivity assumption holds across hardware, featuremetric models could serve as compact, interpretable summaries for comparing devices, with residual variance as a measure of how idiosyncratic circuit behavior is.","A critical test not in the paper is to train on one circuit family and predict another with the same feature values; success would show genuine feature sufficiency, while failure would bound the framework's scope."],"forward_implications":["Volumetric benchmarking plots can be generated from substantially fewer circuits than an exhaustive grid requires, cutting experimental cost while preserving accuracy for interpolation.","Adding two-qubit gate density as a third feature markedly improves predictive accuracy; at one operating point the mean success probability changed from roughly 51 percent to 7 percent as the density varied.","Capability models can predict performance at feature values that were never run, enabling continuous volumetric plots rather than discrete grids.","The same analysis pipeline applies to any capability function and any circuit family, including algorithm-derived families with pseudo-features such as QAOA layer count."],"supporting_citations":[{"why":"Defines the volumetric benchmarking framework that featuremetric benchmarking generalizes.","marker":"[21]"},{"why":"Provides the randomized mirror circuit family and the capability-based methodology used in the demonstrations.","marker":"[22]"},{"why":"Supplies the Gaussian process regression machinery used to build capability models.","marker":"[57]"},{"why":"Introduces monotonic Gaussian process regression used to enforce expected fidelity decay.","marker":"[58]"},{"why":"Formalizes capability learning, the framework within which featuremetric benchmarking is stated.","marker":"[59]"},{"why":"Provides SPAM-error-robust direct fidelity estimation, the method used to estimate process fidelity on the IonQ system.","marker":"[66]"},{"why":"Supplies the Sobol sequence used for quasirandom sampling of feature vectors.","marker":"[75]"}],"fun_headline_variants":["Three circuit features map quantum performance","Few circuits, full picture via featuremetric benchmarks","Gaussian process predicts quantum computer capability","Benchmark reads width, depth, and gate density","Performance maps from sparse circuit data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework assumes that a small set of circuit features explains most of the variation in how accurately a quantum computer runs circuits, so that interpolating between measured circuits predicts untested ones.","fun_headline_variants_meta":{"raw":{"variants":["Three circuit features map quantum performance","Few circuits, full picture via featuremetric benchmarks","Gaussian process predicts quantum computer capability","Benchmark reads width, depth, and gate density","Performance maps from sparse circuit data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000208,"raw_usage":{"total_tokens":1355,"prompt_tokens":847,"completion_tokens":508,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":463,"completion_tokens_details":{"reasoning_tokens":444}},"tokens_in":463,"tokens_out":508,"duration_ms":5662,"temperature":1.0,"reasoning_tokens":444,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:27:52.298832+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run two structurally different circuit families that cover the same feature values on one device, train the Gaussian process on one family, and evaluate it on the other; if the mean absolute prediction error is no better than predicting the global average performance, the feature set is not predictive and the framework's central assumption fails.","supporting_citations":[{"cited_title":"A volumetric framework for quantum computer benchmarks,","cited_arxiv_id":null,"evidence_quote":"Defines the volumetric benchmarking framework that featuremetric benchmarking generalizes."},{"cited_title":"Measuring the capabili- ties of quantum computers,","cited_arxiv_id":null,"evidence_quote":"Provides the randomized mirror circuit family and the capability-based methodology used in the demonstrations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Gaussian process regression machinery used to build capability models."},{"cited_title":"Gaussian processes with monotonicity information,","cited_arxiv_id":null,"evidence_quote":"Introduces monotonic Gaussian process regression used to enforce expected fidelity decay."},{"cited_title":"Learning a quantum computer’s capability,","cited_arxiv_id":null,"evidence_quote":"Formalizes capability learning, the framework within which featuremetric benchmarking is stated."},{"cited_title":"The distribution of points in a cube and the approxi- mate evaluation of integrals,","cited_arxiv_id":null,"evidence_quote":"Supplies the Sobol sequence used for quasirandom sampling of feature vectors."}],"review_version":1}