Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Modeling Quantum Volume Using Randomized Benchmarking of Room-Temperature NV Center Quantum Registers

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Room-temperature NV register hits quantum volume 8.

desk verdict The measured EPGs are a real contribution; the QV=8 headline is a model prediction, not a measurement, and should be presented as such. read the letter →

arxiv 2412.12959 v1 pith:IEAKZZH6 submitted 2024-12-17 quant-ph

classification quant-ph
keywords quantumvolumerandomizedbenchmarkingNVcenterdiamondspinregisternoisemodelheavyoutputprobabilityall-to-allconnectivityroom-temperaturecomputing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper establishes that a three-nuclear-qubit register based on a nitrogen-vacancy (NV) center in diamond, operated at room temperature, can achieve a quantum volume of 8, the maximum possible for three addressable qubits. The figure comes from a noise model calibrated with randomized benchmarking, not from running the quantum-volume circuits on the hardware directly. If the model is faithful, the register passes the heavy-output test at width and depth 3 on every qubit pair, putting a room-temperature solid-state platform on par with other small-scale quantum processors. The result matters because quantum volume is a device-agnostic benchmark, and a reliable room-temperature register would make quantum computation more accessible outside cryogenic environments.

What carries the argument

The carrying object is the depolarizing noise model assembled from randomized benchmarking. For every basis gate, the exponential decay $F = A_0 \cdot \alpha^{N} + B_0$ is fitted to the measured survival probability, and the per-gate error is extracted via $EPG(n,\alpha) = \frac{2^n-1}{2^n}(1-\alpha)$. Single-qubit and two-qubit errors are combined with readout and initialization error estimates, treating each logical two-qubit CNOT between nuclear spins as one depolarizing channel even though its physical implementation uses three electron-mediated CNOTs. This model is then inserted into quantum-volume circuits built from layers of random SU(4) unitaries on random qubit pairs, with success judged by a heavy-output probability threshold of $2/3$.

What would settle it

Perform the quantum-volume heavy-output experiment directly on the same register, without the noise-model simulation, and compare the measured probability at $m=3$ against the $2/3$ threshold; if a direct measurement fails, or only passes with the 15 kHz ODMR-shift discard applied to a substantial fraction of runs, the modeled $V_Q = 8$ would be an optimistic prediction rather than a hardware-level result.

Watch

Extended reading notes

Core claim

The paper's central claim is that a room-temperature NV-center register with three strongly coupled nuclear spins ($^{14}$N, $^{13}$C$_1$, $^{13}$C$_2$) and the electron spin as an ancilla reaches a quantum volume of $V_Q = 2^3 = 8$. This is the outcome of simulating quantum-volume circuits on a noise model calibrated by randomized benchmarking: the circuits of width and depth $m=3$ succeed on all three nuclear-pair connections, keeping the simulated heavy-output probability above the $2/3$ threshold with its full $2\sigma$ spread. The register therefore passes the maximum quantum-volume test possible for a three-qubit system, and the paper takes this as evidence that room-temperature NV centers are a viable platform for quantum computation, with the electron spin's $T_1$ and $T_2$ times acting as the dominant error sources.

Load-bearing premise

The result depends on the assumption that the error rates measured in randomized benchmarking transfer unchanged to the longer, differently structured quantum-volume circuits, and that discarding runs with an ODMR frequency shift above 15 kHz does not remove a failure that would appear in a direct test.

Editorial extensions

If this is right

  • At quantum volume 8, the register reaches the maximum depth-3 heavy-output success allowed by its three-qubit size; larger values are impossible without adding qubits or improving the electron spin's coherence.
  • The calibrated noise model reproduces both the benchmarking decays and the quantum-volume simulations, so it can serve as a predictive tool for other circuit families on this register without additional experimental runs.
  • All-to-all nuclear-spin connectivity means any qubit pair can be entangled through the electron ancilla in a fixed number of steps, avoiding the nearest-neighbor routing overhead common on chip-based devices.
  • Because the electron's $T_1$ and $T_2$ are the stated limiting factors, operating the same register at cryogenic temperatures should push the quantum volume beyond 8.
  • The weakly coupled carbon bath cannot raise the quantum volume even with perfect additional qubits, since the electron relaxation still bounds every two-qubit gate; these spins remain candidates for quantum memory instead.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct heavy-output experiment on this hardware could disagree with the modeled $V_Q=8$ if the depolarizing model misses non-Markovian or context-dependent errors, and the paper provides no direct check of that transfer.
  • The 15 kHz ODMR frequency-shift discard rule is a form of post-selection; reporting the fraction of runs discarded would quantify how much of the quantum-volume result depends on that cut.
  • The same recipe — randomized benchmarking into a depolarizing model, then quantum-volume simulation — could be used to predict quantum volume for other central-spin or solid-state registers; if it generalizes, it becomes a cheap design tool for comparing architectures.
  • For three qubits, quantum volume saturates at 8, so a more informative cross-platform comparison would include the physical overhead behind that number, such as total sequence duration and the number of microwave and radio-frequency pulses.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reports randomized benchmarking of a room-temperature NV-center register with three nuclear-spin qubits (14N, 13C1, 13C2) connected through an electron-spin ancilla. Single- and two-qubit error-per-gate (EPG) values are extracted from exponential decay fits, an error map is constructed, and a depolarizing noise model is built in Qiskit. The authors then simulate quantum-volume (QV) circuits for widths and depths up to m = 3, obtaining a heavy-output probability above the 2/3 threshold on all qubit connections and concluding that the register has a quantum volume of 8. The entire QV result is obtained from simulation with a noise model calibrated on the RB data, not from a direct heavy-output experiment on the hardware.

Significance. If the reported QV = 8 is taken as a validated experimental result, it would be a useful benchmark for room-temperature NV platforms and a clean demonstration of using RB-derived EPGs to predict a composite performance metric. The manuscript has concrete strengths: standard RB protocols with explicit EPG extraction, an all-to-all connectivity map, a reproducible-looking noise-model construction, and a forward simulation of QV circuits. The forward prediction is not definitionally circular, since the QV circuits are not used to fit the model. However, the paper's central claim is substantially weaker than its wording suggests: the QV value is a model prediction with no direct experimental validation, and the model is validated only on the RB data used to construct it. That gap is load-bearing because QV is, by definition, a measured circuit-level property.

major comments (4)
  1. [Section III.C, Fig. 4(c)] The central claim 'The quantum volume circuit for m = 3 was successful on all possible qubit connections, achieving a quantum volume of 8' is based on simulation, not on a direct heavy-output experiment on the register. The text in III.C states 'Our simulations for quantum volume circuits are performed with the noisy model' and Fig. 4(c) labels the heavy-output probability as 'simulated'. Quantum volume is defined by a threshold on measured heavy-output probabilities, so the manuscript should state clearly that QV = 8 is a modeled estimate. A direct m = 3 heavy-output experiment, even with modest statistics, would make the claim experimentally grounded and should be reported if available.
  2. [Section III.B, noise model validation] The consistency check that the noise model 'reproduces the measurement results of the nuclear spins' is performed on the same RB datasets from which the EPG parameters were extracted. This is an in-sample fit and does not test whether the depolarizing model transfers to longer, differently structured QV circuits that involve all three nuclear qubits simultaneously. An out-of-sample check, such as predicting RB decay at sequence lengths beyond the fitted range, or predicting a separate interleaved RB or three-qubit RB measurement, would materially support the extrapolation to QV circuits.
  3. [Section III.A, composite two-qubit gate model] Each logical two-qubit gate between two nuclear spins is implemented by three electron-mediated CNOT gates, but the noise model assigns a single depolarizing error rate to the composite logical gate. The two-qubit RB calibration is performed on the pair under test only, so state-dependent crosstalk, leakage, or heating that appears only when all three nuclear registers are simultaneously active is invisible to the calibration. The authors should either provide evidence that the composite-gate error is independent of the state of the third nuclear spin or use a three-qubit benchmarking protocol to validate the effective two-qubit error channel.
  4. [Section III.C and Section IV, 15 kHz ODMR shift discard rule] The text states that 'Measurements exhibiting a shift in the ODMR frequency of 15 kHz will be discarded' and, earlier, identifies heating-induced frequency shifts as a real error source for 13C2. Fitting the RB EPGs on this post-selected subset means the noise model describes the hardware only under a favorable operating condition. The simulated QV = 8 therefore corresponds to a post-selected (drift-free) register, not to the raw hardware as operated. The authors should quantify the bias introduced by this discard rule, for example by reporting the fraction of discarded runs and the EPG values without the cutoff, or by running a direct QV experiment without this post-selection.
minor comments (5)
  1. [Section II] There are small typographical errors: 'will be refereed to as' should read 'will be referred to as', and the phrase 'will be called c 13C2' contains a stray 'c' before '13C2'.
  2. [Reference [26]] Reference [26] ('Quantum volume. Nature, 2017') is incomplete and does not give authors, volume, or page numbers; please provide a complete citation or use the arXiv/preprint version of the quantum-volume proposal.
  3. [Equation (4)] In Eq. (4), the parameters f1 and f2 are described as 'Azz/2' but no units are given; please specify whether they are frequencies in Hz and how they relate to the hyperfine couplings in Table I.
  4. [Table I] The T2* entry for the electron spin is given as 12.85 µs without an uncertainty, while other entries have uncertainties; adding an error bar and a brief note on how T2* was extracted would improve precision.
  5. [Section III.C] The phrase 'record quantum volume metric for ambient condition operation to date' is not supported by a comparison with other room-temperature platforms; either add a comparative table or soften the claim.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: QV=8 is a forward simulation from independently measured RB error rates, not a restatement of the fit.

full rationale

The paper's derivation chain is: (1) measure single- and two-qubit randomized benchmarking decays and extract error-per-gate via EPG(n, alpha) = (2^n - 1)/2^n * (1 - alpha); (2) assemble a depolarizing noise model using those EPGs plus readout and initialization errors, following the standard construction cited as [19,20]; (3) simulate quantum-volume circuits through that model and count heavy outputs; (4) report QV=8 when the simulated heavy-output probability exceeds 2/3. Step (3) is a forward prediction: the QV heavy-output statistic is not fitted to any QV data, and no equation in the paper re-expresses QV as the RB fit by construction. The consistency statement that the noise model 'reproduces the measurement results' validates the model on the same RB data that produced the EPGs, which is a self-consistency check rather than independent validation, but it is not a circular derivation. The 15 kHz ODMR-shift discard rule is an acknowledged data-selection limitation that bears on whether the simulated QV would match an unfiltered direct measurement, not on whether the QV result reduces to the model inputs. The only self-citation of note is [19] (Finsterhoelzl and Burkard, both authors of the present work) for the noise-model construction; that citation supplies a standard simulation recipe, while the numerical EPG inputs are measured in this paper, so the self-citation is not load-bearing in the sense of importing an unverified uniqueness theorem or ansatz. No step qualifies under the enumerated circularity patterns; the appropriate verdict is no significant circularity.

Assumptions & free parameters 7 free parameters · 4 assumptions · 0 invented entities

The central QV estimate depends on fitted error rates (EPGs) and the assumption that those rates, plus readout and initialization errors, form a complete noise model for arbitrary circuits. No independent QV measurement or external benchmark is provided. The 15 kHz discard rule is a post hoc data selection that makes the model optimistic.

free parameters (7)
  • single-qubit EPG for 14N = (4.4 +/- 0.2) x 10^-3
    Fitted from single-qubit randomized benchmarking decay; used in the noise model.
  • single-qubit EPG for 13C1 = (1.6 +/- 0.3) x 10^-3
    Fitted from RB decay; used in the noise model.
  • single-qubit EPG for 13C2 = (1.0 +/- 0.5) x 10^-3
    Fitted from RB decay; used in the noise model.
  • two-qubit EPG for 14N-13C1 = (2.3 +/- 0.1) x 10^-2
    Fitted from two-qubit RB between 14N and 13C1; used in the QV simulation.
  • two-qubit EPG for 14N-13C2 = (4.7 +/- 0.4) x 10^-2
    Fitted from two-qubit RB between 14N and 13C2; used in the QV simulation.
  • two-qubit EPG for 13C1-13C2 = (2.4 +/- 0.4) x 10^-2
    Fitted from two-qubit RB between 13C1 and 13C2; used in the QV simulation.
  • readout and initialization error rates = not reported numerically
    Mentioned as part of the noise model and calibrated from single-shot readout histograms and post-selection, but values are not stated.
assumptions (4)
  • standard math Randomized benchmarking extracts an average error per gate that is independent of state preparation and measurement errors.
    Standard RB assumption used in Eq. (1) and Eq. (2).
  • domain assumption The two-qubit error per gate measured by RB is a Markovian, gate-independent depolarizing error for the logical CNOT between nuclear spins, despite the gate being implemented with three electron-mediated CNOTs.
    The noise model is built using EPGs as local depolarizing errors; no validation against context-dependent or non-Markovian noise is provided.
  • domain assumption The calibrated noise model remains valid for the longer and differently structured quantum-volume circuits.
    The QV simulation uses the same EPGs and readout/init errors; no QV circuits were run on the hardware.
  • ad hoc to paper Discarding runs with ODMR frequency shift above 15 kHz does not bias the error model.
    The data-exclusion rule removes heating-induced drift, which the paper identifies as a main error source for 13C2; the fraction of discarded runs is not quantified.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Modeling Quantum Volume Using Randomized Benchmarking of Room-Temperature NV Center Quantum Registers." pith.science (2026). https://pith.science/paper/IEAKZZH6

@misc{pith2026241212959,
  author       = {Pith},
  title        = {Pith review of: Modeling Quantum Volume Using Randomized Benchmarking of Room-Temperature NV Center Quantum Registers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IEAKZZH6}},
  note         = {Machine review of arXiv:2412.12959}
}
read the original abstract

Accurately estimating the performance of quantum hardware is crucial for comparing different platforms and predicting the performance and feasibility of quantum algorithms and applications. In this paper, we tackle the problem of benchmarking a quantum register based on the NV center in diamond operating at room temperature. We define the connectivity map as well as single qubit performance. Thanks to an all-to-all connectivity the 2 and 3 qubit gates performance is promising and competitive among other platforms. We experimentally calibrate the error model for the register and use it to estimate the quantum volume, a metric used for quantifying the quantum computational capabilities of the register, of 8. Our results pave the way towards the unification of different architectures of quantum hardware and evaluation of the joint metrics.

Figures

Figures reproduced from arXiv: 2412.12959 by the authors.

Figure 1
Figure 1. FIG. 1 [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2 [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3 [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: FIG. 4 [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: FIG. 5 [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-shot readout error benchmark of the nitrogen-vacancy center's electronic qubit

    cond-mat.mes-hall 2025-05 accept novelty 4.0 of 10

    The multi-shot readout error of an NV center qubit scales as Δ/√N, with a fundamental lower bound of sqrt(2/3), and the authors compute Δ for realistic readout conditions.

Reference graph

Works this paper leans on

30 extracted references · 30 canonical work pages · cited by 1 Pith paper

  1. [1]

    Almudena Carrera Vazquez, Caroline Tornow, Diego Riste, Stefan Woerner, Maika Takita, and Daniel J. Eg- ger. Scaling quantum computing with dynamic circuits, 2024

  2. [2]

    A quantum processor based on co- herent transport of entangled atom arrays

    Dolev Bluvstein, Harry Levine, Giulia Semeghini, Tout T Wang, Sepehr Ebadi, Marcin Kalinowski, Alexander Keesling, Nishad Maskara, Hannes Pichler, Markus Greiner, et al. A quantum processor based on co- herent transport of entangled atom arrays. Nature, 604(7906):451–456, 2022

  3. [3]

    Nature, 614(7949):676–681, 2023

    Suppressing quantum errors by scaling a surface code log- ical qubit. Nature, 614(7949):676–681, 2023

  4. [4]

    Supercon- ducting circuits for quantum information: an outlook

    Michel H Devoret and Robert J Schoelkopf. Supercon- ducting circuits for quantum information: an outlook. Science, 339(6124):1169–1174, 2013

  5. [5]

    Quantum dynamics of single trapped ions

    Dietrich Leibfried, Rainer Blatt, Christopher Monroe, and David Wineland. Quantum dynamics of single trapped ions. Reviews of Modern Physics, 75(1):281, 2003

  6. [6]

    Quantum computa- tions with cold trapped ions

    Juan I Cirac and Peter Zoller. Quantum computa- tions with cold trapped ions. Physical review letters, 74(20):4091, 1995

  7. [7]

    Ladd, Andrew Pan, John M

    Guido Burkard, Thaddeus D. Ladd, Andrew Pan, John M. Nichol, and Jason R. Petta. Semiconductor spin qubits. Rev. Mod. Phys., 95:025003, Jun 2023

  8. [8]

    Spin readout and addressability of phosphorus-donor clusters in silicon

    H B¨ uch, S Mahapatra, Rajib Rahman, Andrea Morello, and MY Simmons. Spin readout and addressability of phosphorus-donor clusters in silicon. Nature communica- tions, 4(1):2017, 2013

Show all 30 references
  1. [9]

    Noise correlations in a 1d silicon spin qubit array.arXiv preprint arXiv:2405.03763, 2024

    MB Donnelly, J Rowlands, L Kranz, YL Hsueh, Y Chung, A V Timofeev, H Geng, P Singh-Gregory, SK Gorman, JG Keizer, et al. Noise correlations in a 1d silicon spin qubit array.arXiv preprint arXiv:2405.03763, 2024

  2. [10]

    Map- ping a 50-spin-qubit network through correlated sensing

    GL Van de Stolpe, DP Kwiatkowski, CE Bradley, J Randall, MH Abobeih, SA Breitweiser, LC Bassett, M Markham, DJ Twitchen, and TH Taminiau. Map- ping a 50-spin-qubit network through correlated sensing. Nature Communications, 15(1):2006, 2024

  3. [11]

    Photonic quantum technologies

    Jeremy L O’brien, Akira Furusawa, and Jelena Vuˇ ckovi´ c. Photonic quantum technologies. Nature Photonics , 3(12):687–695, 2009

  4. [12]

    Knill, D

    E. Knill, D. Leibfried, R. Reichle, J. Britton, R. B. Blakestad, J. D. Jost, C. Langer, R. Ozeri, S. Seidelin, and D. J. Wineland. Randomized benchmarking of quan- tum gates. Phys. Rev. A, 77:012307, Jan 2008

  5. [14]

    Towards a room temperature solid state quantum processor - the nitrogen-vacancy center in diamond

    Philipp Neumann. Towards a room temperature solid state quantum processor - the nitrogen-vacancy center in diamond. PhD thesis, University of Stuttgart, 2012

  6. [15]

    On readout and initialisation fidelity by fi- nite demolition single shot readout

    Majid Zahedian, Max Keller, Minsik Kwon, Javid Javadzade, Jonas Meinel, Vadim Vorobyov, and J¨ org Wrachtrup. On readout and initialisation fidelity by fi- nite demolition single shot readout. Quantum Science and Technology, 9(1):015023, dec 2023

  7. [16]

    Gottesman edited by J

    D. Gottesman edited by J. Samuel and J. Lomonaco.Pro- ceedings of Symposia in Applied Mathematics. American Physical Society, 2010

  8. [17]

    Easwar Magesan, J. M. Gambetta, and Joseph Emerson. Scalable and robust randomized benchmarking of quan- tum processes. Phys. Rev. Lett., 106:180504, May 2011

  9. [18]

    Jarmola, V

    A. Jarmola, V. M. Acosta, K. Jensen, S. Chemerisov, and D. Budker. Temperature- and magnetic-field-dependent longitudinal spin relaxation in nitrogen-vacancy ensem- bles in diamond. Phys. Rev. Lett., 108:197601, May 2012

  10. [19]

    Benchmark- ing quantum error-correcting codes on quasi-linear and central-spin processors

    Regina Finsterhoelzl and Guido Burkard. Benchmark- ing quantum error-correcting codes on quasi-linear and central-spin processors. Quantum Science and Technol- ogy, 8(1):015013, nov 2022

  11. [20]

    Wood, Jake Lishman, Julien Gacon, Si- mon Martiel, Paul D

    Ali Javadi-Abhari, Matthew Treinish, Kevin Krsulich, Christopher J. Wood, Jake Lishman, Julien Gacon, Si- mon Martiel, Paul D. Nation, Lev S. Bishop, Andrew W. Cross, Blake R. Johnson, and Jay M. Gambetta. Quan- tum computing with Qiskit, 2024

  12. [21]

    Janet Anders, Daniel K. L. Oi, Elham Kashefi, Dan E. Browne, and Erika Andersson. Ancilla-driven universal quantum computation. Phys. Rev. A, 82:020301, Aug 2010

  13. [22]

    Quantum networks based on color centers in diamond

    Maximilian Ruf, Noel H Wan, Hyeongrak Choi, Dirk En- glund, and Ronald Hanson. Quantum networks based on color centers in diamond. Journal of Applied Physics, 130(7), 2021. 10

  14. [23]

    Quantum network utility: A framework for bench- marking quantum networks

    Yuan Lee, Wenhan Dai, Don Towsley, and Dirk En- glund. Quantum network utility: A framework for bench- marking quantum networks. Proceedings of the National Academy of Sciences, 121(17):e2314103121, 2024

  15. [24]

    Huang W., Yang C.H., and Chan K.W. et al. Fidelity benchmarks for two-qubit gates in silicon. Nature, 2019

  16. [25]

    Boixo S.and Isakov S.V.and Smelyanskiy V.N. et al. Characterizing quantum supremacy in near-term devices. Nature Physics, 2018

  17. [26]

    L. S. Bishop, S. Bravyi, A. Cross, J. M. Gambetta, , and J. A. Smolin. Quantum volume. Nature, 2017

  18. [27]

    Quantum optimization using varia- tional algorithms on near-term quantum devices

    Nikolaj Moll et al. Quantum optimization using varia- tional algorithms on near-term quantum devices. Quan- tum Science and Technology, 3(3):030503, jun 2018

  19. [28]

    Complexity-theoretic foundations of quantum supremacy experiments, 2016

    Scott Aaronson and Lijie Chen. Complexity-theoretic foundations of quantum supremacy experiments, 2016

  20. [29]

    Validating quantum computers using randomized model circuits

    Cross Andrew W., Bishop Lev S., Sheldon Sarah, Na- tion Paul D., and Gambetta Jay M. Validating quantum computers using randomized model circuits. Phys. Rev. A, 100:032328, Sep 2019

  21. [30]

    A single electron sensor assisted by a quantum coprocessor

    Sebastian Zaiser. A single electron sensor assisted by a quantum coprocessor. PhD thesis, University of Stuttgart, 2019

  22. [31]

    Spin polarization in single spin experiments on defects in diamond

    Iulian Popa, Thorsten Gaebel, Philipp Neumann, Fedor Jelezko, and Joerg Wrachtrup. Spin polarization in single spin experiments on defects in diamond. Israel Journal of Chemistry, 46(4):393–398, 2006

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.