Pith. sign in

REVIEW 4 major objections 6 minor 54 references

A Heterogeneous Distributed Architecture for Quantum Simulation

T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A heterogeneous distributed architecture with a magic core and 1D cold-storage lanes can run fault-tolerant fermionic simulations at speeds comparable to homogeneous designs with several times more magic factories.

desk verdict A genuinely new lane-based architecture with a serious, well-structured evaluation whose exact speedups are provisional until the simulator ships and the 2D baseline is layout-swept. read the letter →

arxiv 2608.09215 v1 pith:T6WSRK54 submitted 2026-08-10 quant-ph

classification quant-ph PACS 03.67.Lx03.67.Mn03.67.-a
keywords heterogeneousdistributedquantumarchitecturefault-tolerantsimulationPaulistringparityFermi-HubbardmodelsparseSachdev-Ye-Kitaevmagicstatefactoriescold-storagelanesBravyi-Kitaevtransform
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is trying to establish that a distributed fault-tolerant quantum computer does not need richly connected, homogeneous modules to run useful fermionic simulations. It proposes a heterogeneous machine: a centralized magic core—containing T-state factories, hot storage, and Bell-pair caches—attached to one-dimensional lanes of cheap cold-storage modules, and argues that this organization matches fermionic workloads, where the parity of a Pauli string (a product of single-qubit Pauli operators) is accumulated with Clifford operations across many qubits while non-Clifford resources act on a single qubit. The central quantitative claim is that a six-lane system with 30 T factories executes a single Trotter step of a 450-logical-qubit Fermi-Hubbard simulation in about 49 s, within roughly 1.4× of a homogeneous 2D distributed device with about 125 T factories and far more connectivity, and about 2× faster when T-factory counts are matched. If true, scaling fault-tolerant simulators becomes a matter of adding low-cost storage lanes rather than replicating expensive magic-generation and compute nodes.

What carries the argument

The load-bearing object is the heterogeneous 1D lane architecture: a magic core holding T-state factories, hot storage, and Bell-pair caches, connected to several one-dimensional chains of cold-storage modules that hold the logical data. The carrying mechanism is lane-wise parity accumulation—local CNOTs compute partial parities inside every occupied lane simultaneously, and a single non-local CNOT per lane feeds those parities to the core, giving parallel random access to Pauli-string parities at a cost almost independent of where the support lies. This same mechanism supplies intra-string parallelism: even when a single high-weight, overlapping Pauli string is the only task running, parity computation is split across all six lanes. Scheduling is a greedy cycle-level policy: commuting groups formed by the Hamiltonian are executed sequentially, and within a group, strings whose qubit and lane sets are disjoint launch concurrently.

What would settle it

Run a cycle-resolved simulation or small-device experiment comparing the six-lane 1D device against the homogeneous 2D grid at 450 logical qubits under measured neutral-atom parameters—actual movement acceleration, serial or parallel Bell-pair generation, and crosstalk—and compare wall-clock times for one Fermi-Hubbard Trotter step under the Bravyi-Kitaev encoding at matched T-factory counts; the central claim fails if the 1D device is not about 2× faster and the gap to the 125-factory 2D device is not near 1.4×.

Watch

Extended reading notes

Core claim

The central discovery is that the dominant cost of executing non-local Pauli strings can be restructured by architecture. In the proposed design the support of each Pauli string is partitioned by storage lane; qubits within each lane accumulate parity locally, and only one non-local CNOT per occupied lane transfers that parity into the central magic core, so a string spread across six lanes costs six non-local operations rather than $O(w)$ hops across its weight-$w$ support. Across Fermi-Hubbard and sparse SYK workloads up to 450 logical qubits, the paper reports that this lane organization is competitive with, and for matched factory counts faster than, a homogeneous 2D grid of identical modules, with the advantage growing as Pauli strings become more non-local and overlapping. The paper also establishes a complementary regime result: when Bell-pair generation is slow (10 ms), extra T factories beyond a few columns add little wall-clock benefit because communication, not magic supply, becomes the bottleneck.

Load-bearing premise

The comparisons stand on the movement and scheduling model, which assumes logical atoms move with constant acceleration of 5500 m/s², cached Bell pairs are always available, entanglement swapping through intermediate modules is free of extra error, and a greedy cycle-level scheduler sees no contention beyond lanes and qubit supports; if real devices move slower or suffer crosstalk, the reported ratios could shift by a large factor.

Editorial extensions

If this is right

  • Fault-tolerant fermionic simulation can run on a modest device: a six-lane system with 30 T factories completes a 450-qubit Fermi-Hubbard Trotter step in about 49 s.
  • The Bravyi-Kitaev fermion-to-qubit encoding's non-local Pauli strings become the favorable choice as system size grows, because lane-parallel parity accumulation makes their irregular, distant supports cheaper to access than the long strings of the Jordan-Wigner encoding.
  • Scaling up is achieved by appending low-cost cold-storage lanes, while the number of magic sites stays fixed at the lane count instead of growing with every module.
  • For workloads dominated by overlapping non-local strings, such as sparse SYK, intra-string parallelism across lanes matters more than adding factories or magic sites.
  • In a slow-communication regime, hardware effort should go into Bell-pair generation rather than additional T factories, since extra factories beyond roughly three columns produce little speedup.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If lane-parity accumulation is as fast as modeled, the same mechanism should accelerate any parity-heavy circuit—stabilizer measurements, parity checks, or state preparation—not just fermionic Trotter steps; this is a testable extension the paper does not make.
  • The paper's reliance on a number of cached Bell pairs equal to the square of the code distance per non-local CNOT means realistic infidelities that force distillation would raise communication costs, potentially shifting the crossover back toward more local encodings.
  • Pairing the storage lanes with qLDPC codes, a direction the paper names for future work, should cut the per-lane physical footprint substantially because the parity-access mechanism is code-agnostic.
  • The heatmap of wall-clock time versus Bell-pair rate and factory count suggests a design rule: size the magic core to the achievable communication rate, not to peak T demand, which could make the architecture auto-tunable per workload.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes a heterogeneous distributed quantum architecture consisting of a central magic core connected to one-dimensional lanes of cold-storage modules, and evaluates it with a custom cycle-level simulator on fault-tolerant Trotterized Fermi-Hubbard and sparse SYK workloads with 98 to 450 logical qubits. The architecture executes Pauli-string parities lane-locally and feeds them to a centralized rotation stage, while comparisons are made against a monolithic baseline (the lane architecture with instantaneous Bell pairs) and a homogeneous 2D grid of identical modules. The headline quantitative claims are that for a 450-qubit Fermi-Hubbard Trotter step a six-lane, 30-T-factory configuration is within about 1.4x of the wall-clock of a 2D device with about 125 T factories, and that at approximately matched T-factory counts the lane architecture is about 2x faster. The paper also includes random-string storage benchmarks, commuting-group statistics, and a breakdown of execution time into movement, Bell-state wait, and T-stage components.

Significance. The architecture idea is timely and the workload-architecture matching argument (BK non-locality versus lane parallelism) is well motivated. The paper is unusually transparent about its modeling assumptions (Table I, Section IV), uses standard Qiskit-generated Hamiltonians, and presents a falsifiable performance claim rather than a qualitative design proposal. The random-string experiments in Fig. 6 and the Pareto framing in Fig. 3 are useful contributions. If the performance comparison survives a sensitivity analysis, the paper would make a solid contribution to distributed fault-tolerant architecture design. However, the quantitative conclusions currently rest on an unreleased simulator and on a single unoptimized 2D baseline, so the significance is conditional.

major comments (4)
  1. [IV-B.3, Fig. 5] The central 1.4x and ~2x comparisons are measured against one particular 2D baseline: a near-square grid with a snake-order qubit-to-module assignment and a fixed routing schedule. No sensitivity analysis is reported for this layout, and the text's only justification for the routing choice is that other routing methods increased the number of hops, which does not address the logical-qubit assignment. Since the greedy scheduler in Section IV-G executes strings with disjoint module supports concurrently, a different assignment can change both the number of non-local hops and the amount of string-level parallelism; given that Bell-pair wait and scheduling parallelism are co-dominant in Figs. 7-8, an optimized assignment could materially narrow the claimed advantage. Please add a layout sweep, a randomized-layout distribution, or an analytical upper bound on the sensitivity, and report how the headline ratios change.
  2. [V, Fig. 5(a), abstract] The abstract's 'matched T-factory counts' comparison is not actually matched. The 1D 6L5T configuration has 30 T factories, while the 2D 1T configuration has approximately 25 factories, and the two configurations use different encodings (BK for the 1D device, JW for the 2D device). The reported ~2x speedup therefore conflates T-factory count with encoding choice and with the 2D baseline's layout, so it is a 'comparable-resource' comparison rather than a controlled one. Please provide a same-encoding, equal-T-factory comparison, including equal physical-qubit budgets if possible, or revise the abstract and Observation 2 to state the exact configurations being compared.
  3. [Appendix X-C, Table II] The code-distance model d = ceil(d0 + 2 log(p_L/p0)/log r) is introduced without a reference or validation, and the fitted constants (d0 = 11, p0 = 5e-7, threshold 7e-3, r = 1e-3/7e-3) are not justified. These distances enter the timing model in multiple ways: movement time scales through the physical length in Eq. (3), each non-local CNOT consumes d^2 Bell pairs, and the distance estimates in Table II differ between encodings and workloads. An error in this ad hoc fit could therefore change relative timings, not merely absolute ones. Please replace it with a published resource-estimation model or demonstrate that the reported speedups are insensitive over a plausible range of distance parameters.
  4. [IV-B, IV-G, Fig. 8] The movement and scheduling model assumes a constant acceleration of 5500 m/s^2, no decoherence or cross-talk during shuttling, cached Bell pairs with entanglement swapping through intermediate modules, and a greedy cycle-level scheduler whose only contention is on lanes and qubit supports. These assumptions are stated clearly, but they are not tested. In the 10 ms Bell-generation regime, communication components dominate the runtime (Fig. 8), so a material overestimate of the achievable movement or entanglement-swapping rate could change the architecture ranking, not just the absolute latencies. Please add a sensitivity study varying the acceleration, including serial (non-parallel) ebit generation, and modeling lane-traffic contention or cross-talk penalties; a simple conservative variant of the communication model would be sufficient.
minor comments (6)
  1. [IV-B.3] Please clarify how empty modules are handled when N_grid is slightly larger than m in the near-square grid; the snake-order assignment and scheduling behavior for unused modules are not specified.
  2. [Table I] The table lists 'Bell pair cache size n = d^2' but n is not defined near the table, and 'Bell pair parallelism O(d)' should specify whether the factor d refers to the code distance or to another quantity.
  3. [Eq. (3)] The movement time formula tau_move = 2 sqrt(L_phys/a) should be derived or referenced; as written, the factor of 2 and the assumed acceleration profile are unclear.
  4. [Fig. 5 caption] The caption's statement that 'end point ratios are against the fastest device' is ambiguous; please state explicitly which configuration is the reference for each ratio in panels (a) and (b).
  5. [Section V] When the text says the 2D 1T configuration has 'approximately 25 factories in total,' it should state the exact count for the 450-qubit case (the grid has 25 modules) and explicitly acknowledge the difference from the 30 factories in 6L5T.
  6. [Overall] No code or data availability statement is included, even though the results rely on a custom simulator that is not shipped; releasing the simulator and workload-generation scripts would substantially strengthen reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the 1.4x/2x comparisons are simulator outputs under a shared cost model, not fits to the headline.

full rationale

The paper's central claims are outputs of a cycle-level simulator (Sections IV-E and IV-G) that applies the same movement equation (Eq. 3), Bell-pair consumption rule (Section IV-C), and T-stage cost model (Section IV-D) to the 1D lane architecture, the 2D homogeneous grid, and the monolithic ablation. The workloads are generated by standard Qiskit routines from published Fermi-Hubbard and sparse SYK Hamiltonians (Appendix X-A), with JW and BK mappings implemented according to Qiskit conventions (Appendix X-B). No parameter is fitted to produce the reported 1.4x or 2x ratios; those ratios are measured simulation outputs. The 1D and 2D architectures are specified independently in Section IV-B and compared under an identical movement and Bell-pair model, so the headline comparison does not reduce to an equation identity or to a fitted quantity. The paper does include self-citations, notably [40] and [41], for switch-infidelity context and for the transversal-operation cost of one Bell pair per physical CNOT; these are standard, externally checkable physical counts and are not invoked as uniqueness theorems or as fitted constraints. They are not load-bearing in a circular sense because the simulation results would not be forced by those citations alone. The most substantive methodological concern is that the 2D baseline's snake-order layout and routing schedule are justified only by the sentence 'We explored other routing methods but they increased the number of hops between modules' (Section IV-B.3), and no sensitivity analysis is given for the qubit-to-module assignment. This is a missing-support/robustness gap, not a circularity: a different baseline layout could change the magnitude of the claimed advantage, but the derivation itself is not equivalent to its inputs. No circular step is therefore identified.

Assumptions & free parameters 5 free parameters · 4 assumptions · 2 invented entities

The central performance claims rest on five free parameters (T-gates per rotation, Bell pair times, code-distance fit, model constants, and SYK degree d), all of which are set by hand or from prior literature rather than measured here. The architecture itself is an invented entity with no independent experimental evidence.

free parameters (5)
  • Code-distance model parameters d0, p0, threshold, r = d0=11, p0=5e-7, two-qubit error rate 1e-3, threshold 7e-3
    The distance estimate in Appendix X-C uses a phenomenological fit d = ceil(d0 + 2 log(pL/p0)/log(r)). These constants come from a fitting formula and materially affect physical-footprint estimates.
  • T-gates per rotation = 30 = 30
    Fixed to 30 based on ref [7] for 1e-6 rotation error; the cost is treated as constant and heavily influences all timings.
  • Mean-degree d for sparse SYK = 3
    The sparse SYK model includes quartets with probability p = d / C(2N-1,3), choosing d=3. This determines the workload structure but is a problem-generation parameter rather than a fit to data.
  • On-site interaction U = 8 and hopping t = 1 for Fermi-Hubbard = U=8, t=1
    Standard model parameters from Qiskit Nature's default, not fitted here.
  • Bell pair generation time = 1 ms or 10 ms
    Chosen to represent optimistic and pessimistic regimes, not fitted to measurements in the paper.
assumptions (4)
  • domain assumption Surface-code logical operations are modeled via transversal gates with Bell pair consumption of d^2 per non-local CNOT.
    Assumed throughout Section IV-C, with cited support from [41]. If this abstraction is optimistic for a real distributed implementation, the wall-clock comparisons are affected.
  • domain assumption A uniform movement model for neutral atoms with acceleration 5500 m/s^2 and atom spacing 12 micrometer.
    Taken from ref [53], used in Eq. 3 for all architectures. It assumes routing and shuttling can be modeled by this simple kinematic cost.
  • ad hoc to paper The phenomenological code-distance formula in Appendix X-C.
    The paper's distance estimates rely on a logarithmic fit d = d0 + 2 log(pL/p0)/log(r); this is a fitting formula that is not independently derived in this work.
  • standard math First-order Trotter-Suzuki decomposition with r=10 steps for distance estimates and r=1 for timing.
    Used to define the workload (Eq. 2); the resulting approximation error is not evaluated in the timing model, which is appropriate for an architectural study but is an assumption about the target application.
invented entities (2)
  • 1D cold-storage lane architecture with central magic core
    purpose: Provide parallel random access to Pauli string parities with reduced connectivity and fewer T factories.
    The architecture is the paper's proposal, evaluated only through its own simulator. It is not yet a physical device; no experimental handle is provided.
  • Parallel Bell pair generation at rate O(d) per directed pool
    purpose: Enable non-local CNOTs at 1 ms generation time.
    The paper cites ref [29] for parallel generation, but the specific rate and reliability model are treated as architectural assumptions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Heterogeneous Distributed Architecture for Quantum Simulation." pith.science (2026). https://pith.science/paper/T6WSRK54

@misc{pith2026260809215,
  author       = {Pith},
  title        = {Pith review of: A Heterogeneous Distributed Architecture for Quantum Simulation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/T6WSRK54}},
  note         = {Machine review of arXiv:2608.09215}
}
abstract

Architectural specialization and distribution can help scale fault-tolerant quantum computers, but may also introduce substantial overheads from communication, routing, and resource duplication. We introduce a heterogeneous distributed architecture in which a magic core is connected to an extensible storage system composed of one-dimensional lanes of specialized cold-storage nodes. The storage system supports parallel random access to Pauli string parities. This organization is particularly well suited to fermionic quantum simulation, enabling parallel execution of the highly non-local Pauli strings arising from these systems. We evaluate the architecture on fault-tolerant simulations of the dynamics of the Fermi-Hubbard and sparse Sachdev-Ye-Kitaev (SYK) models on systems of up to 450 logical qubits. These workloads exhibit complementary communication structures: Fermi-Hubbard produces a spectrum of interactions from local to non-local shaped by lattice geometry, whereas sparse SYK produces highly non-local and overlapping Pauli operators. For a Trotter step of a 450-logical-qubit Fermi-Hubbard workload, a six-lane system with 30 T-state factories is within approximately $1.4\times$ the wall-clock time of a homogeneous distributed architecture with 4 times as many T-state factories and substantially greater connectivity and sites for injecting magic. For matched T-factory counts, our architecture is $\sim 2\times$ faster.

Figures

Figures reproduced from arXiv: 2608.09215 by the authors.

Figure 1
Figure 1. Simulation workflow, Pauli-string structure and device implementation. (a) Simulation workflow. (b) Co-support and incidence matrices for the Jordan–Wigner (JW, top) and Bravyi–Kitaev (BK, bottom) transformations on a 7 × 7 Fermi￾Hubbard Hamiltonian. The co-support matrices indicate the probability that qubits i and j both occur in the support of a Pauli string. The incidence matrices indicate which qubits occur in … view at source ↗
Figure 2
Figure 2. Quantum architectures. Gradient coloring represents design space for sizing. (a) Monolithic device with column(s) of T factories (T), hot-storage zones (H), connected via ebits (E) to data (D) / cold storage qubits. (b-c) Distributed ho￾mogeneous architectures. Yellow lines indicate node-to-node interactions required to implement distributed Pauli strings. Example internal homogeneous node structure shown. (d) Dis￾t… view at source ↗
Figure 3
Figure 3. Wall-clock time against total T factory count Pareto plot. (a-b) 450-qubit simulations at 1ms O(d) Bell state generation rates. The data shown in these plots correspond to the Bravyi–Kitaev (BK) transformation. 100 200 300 400 Number of qubits 25 50 75 100 125 Total T factory count (a) 2D 1 T/module 2D 5 T/module 1D 6-lane dev. Mono. 20-lane dev. 100 200 300 400 Number of qubits 5 10 15 20 25 Number of magic sites (… view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Number of T factories, magic sites and distributed connections. (a) Total number of T factories across different distributed quantum computing architectures as problem size increases. (b) Total number of magic sites for each architec￾ture. For the 1D and monolithic arc…
Figure 6
Figure 6. Figure 6: Compute time for random and work-load specific Pauli strings on different architectures. (a-b) Random simulations. We fix d = 11 here. Each graph shows the total time taken to execute 1000 strings of weight w where for each Pauli string the supports are uniformly rando…
Figure 8
Figure 8. Figure 8: Breakdown of execution time per Pauli string. (a￾d) Average time per Pauli string with identified blockers for different workloads, Bell state speeds, number of qubits and encoding. The heterogeneous architecture separates computation and storage into distinct subsyste…
Figure 7
Figure 7. Figure 7: Communication time and T factory production Heatmap shows wall clock time as a function of Bell gener￾ation rate and T-factory columns. Star indicates the 1D 6L5T device. Bold represents runtime-ratios against a monolithic version of this device. Non-bold represents ru…
Figure 9
Figure 9. Figure 9: Parallelizable Pauli strings and wall clock time. (a, b) Graphs of the theoretical max for different encodings and actually achieved for combinations of encodings, architec￾tures, and workloads (a: Fermi-Hubbard b: SYK), number of simultaneously executable Pauli string…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

54 extracted references · 22 canonical work pages

  1. [1]

    Fermion lattices can be simulated by same-size qubit lattices with $\mathcal{O}(1)$ interaction overhead

    G. Aigner, B. Klaver, M. Lanthaler, and W. Lechner, “Fermion lattices can be simulated by same-size qubit lattices withO(1)interaction overhead,” 2026. [Online]. Available: https://arxiv.org/abs/2605.12600

  2. [2]

    Hamiltonian-based graph-state ansatz for variational quantum algorithms,

    A. Anand and K. R. Brown, “Hamiltonian-based graph-state ansatz for variational quantum algorithms,”Physical Review A, vol. 111, no. 1, p. 012437, 2025

  3. [3]

    Leveraging commuting groups for an effi- cient variational hamiltonian ansatz,

    A. Anand and K. R. Brown, “Leveraging commuting groups for an effi- cient variational hamiltonian ansatz,”Quantum Science and Technology, vol. 10, no. 4, p. 045009, 2025

  4. [4]

    Simulated quantum computation of molecular energies,

    A. Aspuru-Guzik, A. D. Dutoi, P. J. Love, and M. Head-Gordon, “Simulated quantum computation of molecular energies,”Science, vol. 309, no. 5741, pp. 1704–1707, 2005

  5. [5]

    Assessing requirements to scale to practical quantum advantage,

    M. E. Beverland, P. Murali, M. Troyer, K. M. Svore, T. Hoefler, V . Kliuchnikov, G. H. Low, M. Soeken, A. Sundaram, and A. Vaschillo, “Assessing requirements to scale to practical quantum advantage,”

  6. [6]

    Logical quantum processor based on reconfigurable atom arrays,

    D. Bluvstein, S. J. Evered, A. A. Geim, S. H. Li, H. Zhou, T. Manovitz, S. Ebadi, M. Cain, M. Kalinowski, D. Hangleiter, J. P. Bonilla Ataides, N. Maskara, I. Cong, X. Gao, P. Sales Rodriguez, T. Karolyshyn, G. Semeghini, M. J. Gullans, M. Greiner, V . Vuleti ´c, and M. D. Lukin, “Logical quantum processor based on reconfigurable atom arrays,” Nature, vol...

  7. [7]

    More efficient clifford+t synthesis for small-angle rotations and application to trotterization,

    M. Bothe, C. S ¨underhauf, M. J. Witham, E. T. Campbell, and N. S. Blunt, “More efficient clifford+t synthesis for small-angle rotations and application to trotterization,” 2026. [Online]. Available: https://arxiv.org/abs/2605.31544

  8. [8]

    Fermionic quantum computation,

    S. B. Bravyi and A. Y . Kitaev, “Fermionic quantum computation,”An- nals of Physics, vol. 298, no. 1, pp. 210–226, 2002. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0003491602962548

Show all 54 references
  1. [9]

    Distributed quantum computation over noisy channels,

    J. I. Cirac, A. K. Ekert, S. F. Huelga, and C. Macchiavello, “Distributed quantum computation over noisy channels,”Physical Review A, vol. 59, no. 6, pp. 4249–4254, Jun 1999. [Online]. Available: http://dx.doi.org/10.1103/PhysRevA.59.4249

  2. [10]

    A. M. Dalzell, S. McArdle, M. Berta, P. Bienias, C.-F. Chen, A. Gily ´en, C. T. Hann, M. J. Kastoryano, E. T. Khabiboulline, A. Kubica, G. Salton, S. Wang, and F. G. S. L. Brand ˜ao, Quantum Algorithms: A Survey of Applications and End-to-end Complexities. Cambridge University...

  3. [11]

    Restrictions on transversal encoded quantum gate sets,

    B. Eastin and E. Knill, “Restrictions on transversal encoded quantum gate sets,”Phys. Rev. Lett., vol. 102, p. 110502, Mar 2009. [Online]. Available: https://link.aps.org/doi/10.1103/PhysRevLett.102.110502

  4. [12]

    Quantum many-body simulations on digital quantum computers: State-of-the-art and future challenges,

    B. Fauseweh, “Quantum many-body simulations on digital quantum computers: State-of-the-art and future challenges,”Nature Communications, vol. 15, no. 1, p. 2123, 2024. [Online]. Available: https://doi.org/10.1038/s41467-024-46402-9

  5. [13]

    A new data structure for cumulative frequency tables,

    P. M. Fenwick, “A new data structure for cumulative frequency tables,” Software: Practice and Experience, vol. 24, no. 3, pp. 327–336, 1994. [Online]. Available: https://onlinelibrary.wiley.com/doi/abs/10.1002/spe. 4380240306

  6. [14]

    Architecting distributed quantum computers: Design insights from resource estimation,

    D. Filippov, P. Yang, and P. Murali, “Architecting distributed quantum computers: Design insights from resource estimation,” 2026. [Online]. Available: https://arxiv.org/abs/2508.19160

  7. [15]

    Quantum simulation,

    I. Georgescu, S. Ashhab, and F. Nori, “Quantum simulation,”Reviews of Modern Physics, vol. 86, no. 1, p. 153–185, Mar. 2014. [Online]. Available: http://dx.doi.org/10.1103/RevModPhys.86.153

  8. [16]

    How to factor 2048 bit rsa integers with less than a million noisy qubits,

    C. Gidney, “How to factor 2048 bit rsa integers with less than a million noisy qubits,” 2025. [Online]. Available: https: //arxiv.org/abs/2505.15917

  9. [17]

    Opportunities and challenges in fault-tolerant quantum computation,

    D. Gottesman, “Opportunities and challenges in fault-tolerant quantum computation,”arXiv preprint arXiv:2210.15844, 2022

  10. [18]

    Quantum telecomputation,

    L. K. Grover, “Quantum telecomputation,” 1997. [Online]. Available: https://arxiv.org/abs/quant-ph/9704012

  11. [19]

    Electron correlations in narrow energy bands,

    J. Hubbard, “Electron correlations in narrow energy bands,”Proceedings of the Royal Society of London. A. Mathematical and Physical Sciences, vol. 276, no. 1365, pp. 238–257, Nov 1963. [Online]. Available: https://doi.org/10.1098/rspa.1963.0204

  12. [20]

    Transversal architecture for megaquop-scale quantum simulation with neutral atoms,

    R. Ismail, I.-C. Chen, C. Zhao, R. Weiss, F. Liu, H. Zhou, S.-T. Wang, A. Sornborger, and M. Kornja ˇca, “Transversal architecture for megaquop-scale quantum simulation with neutral atoms,”PRX Quantum, vol. 7, p. 020343, May 2026. [Online]. Available: https://link.aps.org/doi/...

  13. [21]

    Network requirements for distributed quantum computation,

    H. Jacinto, E. Gouzien, and N. Sangouard, “Network requirements for distributed quantum computation,”Phys. Rev. Res., vol. 8, p. 013205, Feb 2026. [Online]. Available: https://link.aps.org/doi/10.1103/v9ln- c4v2

  14. [22]

    Quantum computing with Qiskit,

    A. Javadi-Abhari, M. Treinish, K. Krsulich, C. J. Wood, J. Lishman, J. Gacon, S. Martiel, P. D. Nation, L. S. Bishop, A. W. Cross, B. R. Johnson, and J. M. Gambetta, “Quantum computing with Qiskit,” 2024

  15. [23]

    ¨Uber das paulische ¨aquivalenzverbot,

    P. Jordan and E. Wigner, “ ¨Uber das paulische ¨aquivalenzverbot,” Zeitschrift f ¨ur Physik, vol. 47, no. 9, pp. 631–651, 1928

  16. [24]

    Impact of network constraints on fault- tolerant distributed quantum computing,

    E. Kaur, S. Pouryousef, N. K. Chandra, H. Shapourian, J. Zhao, R. Kompella, and R. Nejabati, “Impact of network constraints on fault- tolerant distributed quantum computing,” 2026. [Online]. Available: https://arxiv.org/abs/2606.17495 12

  17. [25]

    Architecting early fault toler- ant neutral atoms systems with quantum advantage,

    S. Khan, S. Sethi, K. Sahay, Y . Lin, J. Alnas, S. Kurapati, A. Anand, J. M. Baker, and K. R. Brown, “Architecting early fault toler- ant neutral atoms systems with quantum advantage,”arXiv preprint arXiv:2604.19735, 2026

  18. [26]

    Shorter quantum circuits via single-qubit gate approximation,

    V . Kliuchnikov, K. Lauter, R. Minko, A. Paetznick, and C. Petit, “Shorter quantum circuits via single-qubit gate approximation,” Quantum, vol. 7, p. 1208, Dec. 2023. [Online]. Available: http: //dx.doi.org/10.22331/q-2023-12-18-1208

  19. [27]

    Even more efficient quantum computations of chemistry through tensor hypercontraction,

    J. Lee, D. W. Berry, C. Gidney, W. J. Huggins, J. R. McClean, N. Wiebe, and R. Babbush, “Even more efficient quantum computations of chemistry through tensor hypercontraction,”PRX Quantum, vol. 2, p. 030305, Jul 2021. [Online]. Available: https://link.aps.org/doi/10. 1103/PRXQ...

  20. [28]

    Asymptotically optimal depth fermionic permutation on 2d grid quantum architecture without ancillas,

    D. Li, S. Xu, and Y . Ding, “Asymptotically optimal depth fermionic permutation on 2d grid quantum architecture without ancillas,” 2026. [Online]. Available: https://arxiv.org/abs/2605.26041

  21. [29]

    Parallelized telecom quantum networking with an ytterbium-171 atom array,

    L. Li, X. Hu, Z. Jia, W. Huie, W. K. C. Sun, Aakash, Y . Dong, N. Hiri-O-Tuppa, and J. P. Covey, “Parallelized telecom quantum networking with an ytterbium-171 atom array,”Nature Physics, vol. 21, no. 11, pp. 1826–1833, Sep 2025. [Online]. Available: http://dx.doi.org/10.1038/...

  22. [30]

    A game of surface codes: Large-scale quantum computing with lattice surgery,

    D. Litinski, “A game of surface codes: Large-scale quantum computing with lattice surgery,”Quantum, vol. 3, p. 128, 2019

  23. [31]

    Universal quantum simulators,

    S. Lloyd, “Universal quantum simulators,”Science, vol. 273, no. 5278, pp. 1073–1078, 1996. [Online]. Available: https://www.science.org/doi/ abs/10.1126/science.273.5278.1073

  24. [32]

    Fast simulation of fermions with reconfigurable qubits,

    N. Maskara, M. Kalinowski, D. Gonzalez-Cuadra, and M. D. Lukin, “Fast simulation of fermions with reconfigurable qubits,” 2025. [Online]. Available: https://arxiv.org/abs/2509.08898

  25. [33]

    Quantum computational chemistry,

    S. McArdle, S. Endo, A. Aspuru-Guzik, S. C. Benjamin, and X. Yuan, “Quantum computational chemistry,”Rev. Mod. Phys., vol. 92, p. 015003, Mar 2020. [Online]. Available: https://link.aps.org/doi/10.1103/ RevModPhys.92.015003

  26. [34]

    Scaling the ion trap quantum processor,

    C. Monroe and J. Kim, “Scaling the ion trap quantum processor,” Science, vol. 339, no. 6124, pp. 1164–1169, 2013. [Online]. Available: https://www.science.org/doi/abs/10.1126/science.1231298

  27. [35]

    Heterogeneous architectures enable a 138x reduction in physical qubit requirements for fault- tolerant quantum computing under detailed accounting,

    P. S. Mundada, A. Khindanov, Y . Wang, C. L. Edmunds, P. Coote, M. J. Biercuk, Y . Baum, and M. Hush, “Heterogeneous architectures enable a 138x reduction in physical qubit requirements for fault- tolerant quantum computing under detailed accounting,” 2026. [Online]. Available...

  28. [36]

    Distribution complexity of electronic structure simulations on quantum supercomputers,

    J. Necaise, N. Anand, G. Gyawali, K. G. Johnson, J. D. Whitfield, and M. Mohseni, “Distribution complexity of electronic structure simulations on quantum supercomputers,” 2026. [Online]. Available: https://arxiv.org/abs/2606.20805

  29. [37]

    High-fidelity teleportation of a logical qubit using transversal gates and lattice surgery,

    C. Ryan-Anderson, N. C. Brown, C. H. Baldwin, J. M. Dreiling, C. Foltz, J. P. Gaebler, T. M. Gatterman, N. Hewitt, C. Holliman, C. V . Horst, J. Johansen, D. Lucchetti, T. Mengle, M. Matheny, Y . Matsuoka, K. Mayer, M. Mills, S. A. Moses, B. Neyenhuis, J. Pino, P. Siegfried, R...

  30. [38]

    Injeqt: Improved magic-state injection protocol for fault-tolerant quantum ex- tractor architectures,

    S. Sethi, S. Khan, A. Awasthi, A. Anand, and J. M. Baker, “Injeqt: Improved magic-state injection protocol for fault-tolerant quantum ex- tractor architectures,”arXiv preprint arXiv:2604.25094, 2026

  31. [39]

    Fault-tolerant quantum computation,

    P. W. Shor, “Fault-tolerant quantum computation,” inProceedings of 37th conference on foundations of computer science. IEEE, 1996, pp. 56–65

  32. [40]

    Impact of a free space optics switch on ebit fidelity for distributed quantum computing,

    J. Stack, Y . Liu, A. Ruocco, G. Zervas, and A. Beghelli, “Impact of a free space optics switch on ebit fidelity for distributed quantum computing,” inQuantum Computing, Communication, and Simulation VI, P. R. Hemmer, A. L. Migdall, and I. A. Burenkov, Eds., vol. 13919, Intern...

  33. [41]

    Transversal fault tolerant distributed quantum computing operations,

    J. Stack, M. Wang, and F. Mueller, “Transversal fault tolerant distributed quantum computing operations,”Nature Communications, Jul. 2026. [Online]. Available: http://dx.doi.org/10.1038/s41467-026-75693-3

  34. [42]

    Hetec: Architectures for heterogeneous quantum error correction codes,

    S. Stein, S. Xu, A. W. Cross, T. J. Yoder, A. Javadi-Abhari, C. Liu, K. Liu, Z. Zhou, C. Guinn, Y . Ding, Y . Ding, and A. Li, “Hetec: Architectures for heterogeneous quantum error correction codes,” inProceedings of the 30th ACM International Conference on Architectural Suppo...

  35. [43]

    High-rate, high-fidelity entanglement of qubits across an elementary quantum network,

    L. J. Stephenson, D. P. Nadlinger, B. C. Nichol, S. An, P. Drmota, T. G. Ballance, K. Thirumalai, J. F. Goodwin, D. M. Lucas, and C. J. Ballance, “High-rate, high-fidelity entanglement of qubits across an elementary quantum network,”Phys. Rev. Lett., vol. 124, p. 110501, Mar 2...

  36. [44]

    Generalized trotter’s formula and systematic approximants of exponential operators and inner derivations with applications to many- body problems,

    M. Suzuki, “Generalized trotter’s formula and systematic approximants of exponential operators and inner derivations with applications to many- body problems,”Communications in Mathematical Physics, vol. 51, no. 2, pp. 183–190, 1976

  37. [45]

    Circuit optimization of hamiltonian simulation by simultaneous diagonalization of pauli clusters,

    E. Van Den Berg and K. Temme, “Circuit optimization of hamiltonian simulation by simultaneous diagonalization of pauli clusters,”Quantum, vol. 4, p. 322, 2020

  38. [46]

    Communication links for distributed quantum computation,

    R. Van Meter, K. Nemoto, and W. Munro, “Communication links for distributed quantum computation,”IEEE Transactions on Computers, vol. 56, no. 12, pp. 1643–1653, Dec 2007. [Online]. Available: http://dx.doi.org/10.1109/TC.2007.70775

  39. [47]

    The pinnacle architecture: Reducing the cost of breaking rsa-2048 to 100 000 physical qubits using quantum ldpc codes,

    P. Webster, L. Berent, O. Chandra, E. T. Hockings, N. Baspin, F. Thomsen, S. C. Smith, and L. Z. Cohen, “The pinnacle architecture: Reducing the cost of breaking rsa-2048 to 100 000 physical qubits using quantum ldpc codes,” 2026. [Online]. Available: https://arxiv.org/abs/2602.11457

  40. [48]

    Roofline: an insightful visual performance model for multicore architectures,

    S. Williams, A. Waterman, and D. Patterson, “Roofline: an insightful visual performance model for multicore architectures,”Communications of the ACM, vol. 52, no. 4, pp. 65–76, 2009

  41. [49]

    A sparse model of quantum holography,

    S. Xu, L. Susskind, Y . Su, and B. Swingle, “A sparse model of quantum holography,” 2020. [Online]. Available: https://arxiv.org/abs/2008.02303

  42. [50]

    Tour de gross: A modular quantum computer based on bivariate bicycle codes,

    T. J. Yoder, E. Schoute, P. Rall, E. Pritchett, J. M. Gambetta, A. W. Cross, M. Carroll, and M. E. Beverland, “Tour de gross: A modular quantum computer based on bivariate bicycle codes,” Jun 2025, arXiv:2506.03094 [quant-ph]. [Online]. Available: http://arxiv.org/abs/2506.03094

  43. [52]

    A universal quantum information preserving photonic switch for scalable quantum networks,

    J. Zhao, S. Vinet, A. Minoofar, M. Kilzer, L. Wang, G. Moody, V . Pandey, R. Kompella, and R. Nejabati, “A universal quantum information preserving photonic switch for scalable quantum networks,”

  44. [53]

    Resource analysis of low-overhead transversal architectures for reconfigurable atom arrays,

    H. Zhou, C. Duckering, C. Zhao, D. Bluvstein, M. Cain, A. Kubica, S.-T. Wang, and M. D. Lukin, “Resource analysis of low-overhead transversal architectures for reconfigurable atom arrays,” inProceedings of the 52nd Annual International Symposium on Computer Architecture, ser. ...

  45. [2022]

    Available: https://arxiv.org/abs/2211.07629

    [Online]. Available: https://arxiv.org/abs/2211.07629

  46. [2026]

    Available: https://arxiv.org/abs/2604.21902

    [Online]. Available: https://arxiv.org/abs/2604.21902

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.