REVIEW 3 major objections 5 minor 2 cited by
TensorQC: Towards Scalable Distributed Quantum Computing via Tensor Networks
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read TensorQC claims that using tensor network contraction for classical post-processing turns circuit cutting's exponential reconstruction cost into an exponential in a smaller quantity, and demonstrates 200-qubit benchmarks on a single GPU.
desk verdict A genuinely useful tensor-network method for circuit-cutting post-processing with a proof flaw that kills the general exponential-advantage claim; referee it with a mandate to fix the theory. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the tensor network formed by the subcircuit outputs: each subcircuit's output is a tensor whose indices are the cut edges it touches, every cut edge has dimension four, one per Pauli basis $\{I, X, Y, Z\}$, and reconstructing the original output is a sequence of pairwise tensor contractions over shared cut edges. The cost identity is that each contraction step costs $4^{\#\text{cuts in that step}}$ multiplications, so the total is $O(4^{K_{\max}} m)$ with $K_{\max}$ the maximum number of cut edges at any one step, replacing the $O(4^{|E|} m)$ brute-force bound. Heavy state selection (HSS) keeps only the highest-L2-norm binary states of each subcircuit before contraction, reducing the output dimension substantially. The greedy graph-growing algorithm merges neighboring circuit fragments by an estimated merging cost that penalizes violations of QPU constraints and large contraction-edge counts, approximating a solution to the NP-complete constrained graph partitioning problem that finds cuts.
What would settle it
Construct a partition with four subcircuits where two of the subcircuits jointly touch every cut edge, so that the first tensor contraction has $K_1 = |E|$; then TensorQC's cost is $O(4^{|E|})$ and the gap to the $O(4^{|E|} m)$ baseline becomes only the factor $m$. A wall-clock comparison on such a network, or a direct count of contractions, would settle whether the exponential advantage survives outside the paper's demonstrated benchmarks.
Extended reading notes
Core claim
The central discovery is that the classical co-processing step of circuit cutting is a tensor network contraction, so it inherits the cost savings of contraction-order optimization. Prior methods evaluate the reconstruction formula by iterating over all $4^{|E|}$ basis permutations of the cut edges and multiplying every subcircuit output from scratch, costing $O(4^{|E|} m)$ multiplications. TensorQC instead contracts subcircuit tensors pairwise along their shared cut edges, reusing intermediate products and exploiting the distributive property of multiplication, with total cost bounded by $O(4^{K_{\max}} m)$; the paper argues that $K_{\max} < |E|$ whenever more than three subcircuits are produced, which is the source of the exponential gap. On top of this, heavy state selection prunes each subcircuit's output to a few significant binary states, cutting the reconstructed dimension from $2^n$ to a product of small per-subcircuit sets, and a greedy graph-growing heuristic searches for cuts under QPU width, gate-count, and contraction-edge constraints. The paper reports end-to-end runtimes for six benchmarks up to 200 qubits on a single GPU, with QPU quantum-area requirements reduced by more than $10\times$.
Load-bearing premise
The exponential speedup rests on the unproven assertion that $K_{\max} < |E|$ whenever a circuit is split into more than three subcircuits, and the large-benchmark demonstrations assume that random subcircuit outputs faithfully reproduce the post-processing cost structure of real QPU data.
Editorial extensions
If this is right
- Distributed hybrid quantum computing becomes practical for circuits that exceed any single QPU's capacity: the paper runs six benchmarks up to 200 qubits with one GPU for classical post-processing.
- QPU resource requirements, measured as the quantum area of the largest subcircuit, shrink by more than $10\times$, relaxing both qubit-count and error-rate demands on near-term hardware.
- For output landscapes with a few dominant states, such as GHZ and W-state circuits, heavy state selection reconstructs nearly all probability mass while sampling far fewer than one part per million of the binary states.
- For distributed output landscapes, such as random supremacy circuits, heavy state selection retains a smaller fraction of the amplitude but still approaches the maximum possible retention for a fixed state budget.
- The classical post-processing cost advantage over the prior parallelized reconstruction method can exceed $10^9\times$ on the demonstrated benchmarks.
Reading between the lines
- The exponential advantage is only guaranteed when no pair of subcircuits collectively touches all cut edges; partitionings in which a small cluster of subcircuits covers all cuts would reduce the speedup to roughly linear, a regime the paper does not analyze.
- Because heavy state selection is most effective on skewed output distributions, the method suggests a practical workload-selection rule: prefer cutting circuits whose amplitude landscape concentrates on few states, such as optimization and state-preparation circuits.
- The paper's multiplication-count cost model is a proxy for wall-clock runtimes; a GPU-profiling-aware cost model would likely change the cut-finding decisions and tighten the runtime estimates, as the paper itself notes.
- Wire cutting along qubit wires is orthogonal to gate-level circuit knitting, so combining both would allow partitioning at arbitrary two-qubit gates and could expand the space of feasible, low-cost cuts.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. TensorQC proposes to replace the brute-force reconstruction step of quantum circuit cutting with tensor network contraction, together with a heavy-state-selection heuristic and a greedy graph-growing algorithm to locate cuts. The paper claims an exponential runtime advantage over prior parallelization (CutQC), formalized as O(4^{Kmax} m) versus O(4^{|E|} m), and reports benchmarks up to 200 qubits run on a single GPU with reduced quantum area requirements.
Significance. The tensor-network equivalence for circuit-cutting reconstruction is mathematically sound, and the cost model provides a useful framework for reasoning about hybrid quantum-classical runtime. The benchmark suite covers diverse circuits, and the reported contraction-time scaling on a GPU is a promising indicator. However, the general exponential-advantage claim is not established, and the experiments use synthetic subcircuit outputs rather than actual QPU execution, so the practical significance as stated is not yet demonstrated.
major comments (3)
- [Section III-D, Eqs. (4)-(5)] The claim that Kmax < |E| whenever there are more than three subcircuits is false. Consider a star topology with m=4 subcircuits: a central subcircuit C1 connected to leaves C2, C3, C4 by three cut edges, so |E|=3. Any contraction sequence that produces the final state must at some step merge the tensor containing C1 with the rest, and at that step all three cut edges are present as open indices; pre-contracting the leaves does not remove them. Hence every contraction order has Kmax=3=|E|, contradicting the proof in Section III-D. The exponential advantage over CutQC therefore does not hold for arbitrary cut networks; it can only be claimed for restricted topologies (e.g., bounded treewidth) or for the specific cut graphs arising in the benchmarks, and the manuscript should state and prove a conditional theorem.
- [Section VII-A and Abstract] The abstract asserts that the benchmarks run "using QPUs available nowadays," but Section VII-A states that QPUs are too small and noisy, and instead the authors "use random numbers as the subcircuit output." The 200-qubit results are therefore classical post-processing demonstrations on synthetic vectors, not end-to-end hybrid executions on QPUs. This also undermines the HSS evaluation in Section VIII-D, because the L2 norms that drive Algorithm 1 are computed from random data rather than from true subcircuit probability distributions, so the reported amplitude-retention ratios do not validate the protocol for real circuit outputs.
- [Section VIII-B, Figure 10] The "exponential classical overhead advantage" is computed as a ratio of multiplication counts between the optimized tensor-network contraction and CutQC's brute-force reconstruction. Since the tensor-network cost is optimized by CoTenGra, the reported ratios depend on the cut graph structure and do not substantiate the general complexity-theoretic claim of Section III-D. The paper should either provide a rigorous bound on Kmax for the benchmark cut graphs or present Figure 10 as an empirical observation for those circuits only, without extrapolating to a universal exponential advantage.
minor comments (5)
- [Section III-C, Eq. (5)] The cost of the prior method for m subcircuits is 4^{|E|} (m-1) scalar multiplications, not 4^{|E|} m; the factor m appears to be an overcount in the asymptotic complexity expression.
- [Section V, Eq. (6)] The QPU runtime model multiplies by 2^{wi} while also stating that the number of shots equals the number of states per subcircuit; the relationship between the shot count and the factor 2^{wi} should be clarified, since a single factor may be double-counting the number of subcircuit executions.
- [Algorithm 1] The loop description in lines 2-6 of Algorithm 1 is ambiguous: the statement "Add arg max ||pj,i|| to xj" does not specify the range of the argmax, and the "Remove max" step is unclear about whether the state is removed from the candidate set or from the subcircuit output.
- [Figures 7-11] Figure 7 includes error bars, but Figures 8-11 do not; the paper should state whether these are single runs or averages and, if averages, why error bars are omitted.
- [Section VII-A] The assumption of ten QPUs is stated but not justified against actual cloud availability, and it is unclear whether the QPU runtime in Eq. (6) is divided by ten or treated as a total across all QPUs.
Circularity Check
No significant circularity: TensorQC's runtime comparison is an algebraic accounting identity with no fitted parameters; the only caveats (Section III-D topology proof gap, random subcircuit outputs) are correctness or limitation issues, not circularity.
full rationale
The central claim is a direct algebraic comparison rather than a fitted or self-referential prediction. Equation (1) defines the brute-force reconstruction as summing over 4^{|E|} edge-base permutations, giving cost O(4^{|E|} m). Equations (3) and (4) express the same summation as tensor contractions whose per-step cost is 4^{#cuts}, bounded by O(4^{Kmax} m). Kmax is obtained from the tensor network topology, not calibrated to measured runtimes. The runtime model in Equations (6)-(8) uses fixed hardware timings (tg = 10^{-7} s, tm = 10^{-6} s), fixed GPU throughput, and hand-set heuristic thresholds (Kt = 10, Qmax = 10^4); no parameter is fitted to the wall-clock points in Figure 7, and the paper explicitly notes its estimates tend to overestimate, which is inconsistent with post-hoc fitting. The only same-author citations ([47], [48]) describe the baseline parallel reconstruction being compared and are not load-bearing for TensorQC's own derivation. Two non-circular concerns are noted: Section VII-A states that large-benchmark subcircuit outputs are random numbers, so the large-scale runtime demonstrations are shape-preserving but not end-to-end QPU executions; and Section III-D's proof that Kmax < |E| for more than three subcircuits is not generally valid (a four-subcircuit cycle can have every first contraction touching all edges). These are correctness or methodology limitations, not circular reductions of the paper's results to their own inputs.
Assumptions & free parameters
free parameters (4)
- Kt =
10
- Qmax =
10^4
- smax (HSS state limit) =
28
- Number of shots per subcircuit =
2^w_i with limits 2^10 to 2^20
assumptions (4)
- standard math The circuit-cutting reconstruction formula (Eq. 1) equals a tensor network contraction over cut-edge indices.
- domain assumption The cost of contracting a tensor network is captured by counting multiplications as 4^#cuts per contraction step, independent of output-state dimensions.
- ad hoc to paper For any partition into more than three subcircuits, an optimal contraction sequence has Kmax < |E|.
- ad hoc to paper The greedy graph-growing heuristic and the threshold table (Table I) approximate the solution to the NP-hard cut-partition problem.
Cite this review
Pith. "Pith review of TensorQC: Towards Scalable Distributed Quantum Computing via Tensor Networks." pith.science (2026). https://pith.science/paper/IPIMPEYJ
@misc{pith2026250203445,
author = {Pith},
title = {Pith review of: TensorQC: Towards Scalable Distributed Quantum Computing via Tensor Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/IPIMPEYJ}},
note = {Machine review of arXiv:2502.03445}
}
abstract
A quantum processing unit (QPU) must contain a large number of high quality qubits to produce accurate results for problems at useful scales. In contrast, most scientific and industry classical computation workloads happen in parallel on distributed systems, which rely on copying data across multiple cores. Unfortunately, copying quantum data is theoretically prohibited due to the quantum non-cloning theory. Instead, quantum circuit cutting techniques cut a large quantum circuit into multiple smaller subcircuits, distribute the subcircuits on parallel QPUs and reconstruct the results with classical computing. Such techniques make distributed hybrid quantum computing (DHQC) a possibility but also introduce an exponential classical co-processing cost in the number of cuts and easily become intractable. This paper presents TensorQC, which leverages classical tensor networks to bring an exponential runtime advantage over state-of-the-art parallelization post-processing techniques. As a result, this paper demonstrates running benchmarks that are otherwise intractable for a standalone QPU and prior circuit cutting techniques. Specifically, this paper runs six realistic benchmarks using QPUs available nowadays and a single GPU, and reduces the QPU size and quality requirements by more than $10\times$ over purely quantum platforms.
Figures
Figures from the paper (9 more)
Forward citations
Cited by 2 Pith papers
-
Towards a Utility-Scale Quantum Edge Detection for Real-World Medical Image Data
A two-level decomposition of images and circuits lets Quantum Hadamard Edge Detection run with 62% lower depth, 93% fewer CNOTs, and 95.6% simulated fidelity, demonstrated on real MRI data.
-
Solving Large-Scale Vehicle Routing Problems with Hybrid Quantum-Classical Decomposition
A standard graph partitioner and circuit-cutting toolkit shrink a 13-node VRP from 156 qubits to 6-qubit subcircuits, but the quality of the 13-node solution is not reported.
Reference graph
Works this paper leans on
-
[47]
Cutting quantum circuits to run on quantum and classical platforms,
W. Tang and M. Martonosi, “Cutting quantum circuits to run on quantum and classical platforms,” arXiv preprint arXiv:2205.05836 , 2022
arXiv 2022
-
[1]
Universal gates for protected superconducting qubits using optimal control,
M. Abdelhafez, B. Baker, A. Gyenis, P. Mundada, A. A. Houck, D. Schuster, and J. Koch, “Universal gates for protected superconducting qubits using optimal control,” Physical Review A , vol. 101, no. 2, p. 022321, 2020
work page 2020
-
[2]
K. Andreev and H. Racke, “Balanced graph partitioning,” Theory of Computing Systems, vol. 39, no. 6, pp. 929–939, 2006
work page 2006
-
[3]
ARQUIN: Architectures for multinode superconducting quantum computers,
J. Ang, G. Carini, Y . Chen, I. Chuang, M. Demarco, S. Economou, A. Eickbusch, A. Faraon, K.-M. Fu, S. Girvin, M. Hatridge, A. Houck, P. Hilaire, K. Krsulich, A. Li, C. Liu, Y . Liu, M. Martonosi, D. C. McKay, J. Misewich, M. Ritter, R. J. Schoelkopf, S. A. Stein, S. Suss- man, H. X. Tang, W. Tang, T. Tomesh, N. M. Tubman, C. Wang, N. Wiebe, Y .-X. Yao, D...
work page 2024
-
[4]
Complexity of finding embeddings in ak-tree,
S. Arnborg, D. G. Corneil, and A. Proskurowski, “Complexity of finding embeddings in ak-tree,” SIAM Journal on Algebraic Discrete Methods , vol. 8, no. 2, pp. 277–284, 1987
work page 1987
-
[5]
Quantum supremacy using a programmable superconducting processor,
F. Arute, K. Arya, R. Babbush, D. Bacon, J. C. Bardin, R. Barends, R. Biswas, S. Boixo, F. G. S. L. Brandao, D. A. Buell, B. Burkett, Y . Chen, Z. Chen, B. Chiaro, R. Collins, W. Courtney, A. Dunsworth, E. Farhi, B. Foxen, A. Fowler, C. Gidney, M. Giustina, R. Graff, K. Guerin, S. Habegger, M. P. Harrigan, M. J. Hartmann, A. Ho, M. Hoffmann, T. Huang, T. ...
2019
-
[6]
Approximate quantum fourier transform and decoherence,
A. Barenco, A. Ekert, K.-A. Suominen, and P. T ¨orm¨a, “Approximate quantum fourier transform and decoherence,”Physical Review A, vol. 54, no. 1, p. 139, 1996
work page 1996
-
[7]
i-qer: An intelligent approach towards quantum error reduction,
S. Basu, A. Saha, A. Chakrabarti, and S. Sur-Kolay, “i-qer: An intelligent approach towards quantum error reduction,” ACM Transactions on Quantum Computing, 2021
work page 2021
Show all 57 references
-
[8]
Characterizing quantum supremacy in near-term devices,
S. Boixo, S. V . Isakov, V . N. Smelyanskiy, R. Babbush, N. Ding, Z. Jiang, M. J. Bremner, J. M. Martinis, and H. Neven, “Characterizing quantum supremacy in near-term devices,”Nature Physics, vol. 14, no. 6, pp. 595–600, 2018
2018
-
[9]
The future of quantum computing with superconducting qubits,
S. Bravyi, O. Dial, J. M. Gambetta, D. Gil, and Z. Nazario, “The future of quantum computing with superconducting qubits,” Journal of Applied Physics, vol. 132, no. 16, 2022
2022
-
[10]
Subsys- tem surface codes with three-qubit check operators,
S. Bravyi, G. Duclos-Cianci, D. Poulin, and M. Suchara, “Subsys- tem surface codes with three-qubit check operators,” arXiv preprint arXiv:1207.1443, 2012
2012 arXiv
-
[11]
On optimizing a class of multi-dimensional loops with reduction for parallel execution,
L. Chi-Chung, P. Sadayappan, and R. Wenger, “On optimizing a class of multi-dimensional loops with reduction for parallel execution,” Parallel Processing Letters, vol. 7, no. 02, pp. 157–168, 1997
1997
-
[12]
Deterministic construction of arbitrary w states with quadratically increasing number of two-qubit gates,
F. Diker, “Deterministic construction of arbitrary w states with quadratically increasing number of two-qubit gates,” arXiv preprint arXiv:1606.09290, 2016
2016 arXiv
-
[13]
Systematic crosstalk mitigation for superconducting qubits via frequency-aware compilation,
Y . Ding, P. Gokhale, S. F. Lin, R. Rines, T. Propson, and F. T. Chong, “Systematic crosstalk mitigation for superconducting qubits via frequency-aware compilation,” in 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 2020, pp. 201–214
2020
-
[14]
Three qubits can be entangled in two inequivalent ways,
W. D ¨ur, G. Vidal, and J. I. Cirac, “Three qubits can be entangled in two inequivalent ways,” Physical Review A, vol. 62, no. 6, p. 062314, 2000
2000
-
[15]
Doubling the size of quantum simulators by entanglement forging,
A. Eddins, M. Motta, T. P. Gujarati, S. Bravyi, A. Mezzacapo, C. Had- field, and S. Sheldon, “Doubling the size of quantum simulators by entanglement forging,” PRX Quantum, vol. 3, no. 1, p. 010309, 2022
2022
-
[16]
A quantum approximate optimization algorithm,
E. Farhi, J. Goldstone, and S. Gutmann, “A quantum approximate optimization algorithm,” arXiv preprint arXiv:1411.4028 , 2014
2014 arXiv
-
[17]
Surface codes: Towards practical large-scale quantum computation,
A. G. Fowler, M. Mariantoni, J. M. Martinis, and A. N. Cleland, “Surface codes: Towards practical large-scale quantum computation,” Physical Review A, vol. 86, no. 3, p. 032324, 2012
2012
-
[18]
Hyper-optimized tensor network contraction,
J. Gray and S. Kourtis, “Hyper-optimized tensor network contraction,” Quantum, vol. 5, p. 410, 2021
2021
-
[19]
Going beyond bell’s theorem,
D. M. Greenberger, M. A. Horne, and A. Zeilinger, “Going beyond bell’s theorem,” in Bell’s theorem, quantum theory and conceptions of the universe. Springer, 1989, pp. 69–72
1989
-
[20]
A fast quantum mechanical algorithm for database search,
L. K. Grover, “A fast quantum mechanical algorithm for database search,” in Proceedings of the twenty-eighth annual ACM symposium on Theory of computing , 1996, pp. 212–219
1996
-
[21]
Gpipe: Efficient training of giant neural networks using pipeline parallelism,
Y . Huang, Y . Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V . Le, Y . Wu, and Z. Chen, “Gpipe: Efficient training of giant neural networks using pipeline parallelism,” Advances in neural information processing systems , vol. 32, 2019
2019
-
[22]
Ibm quantum,
IBM, “Ibm quantum,” 2021, https://quantum.ibm.com/
2021
-
[23]
Optimized surface code communi- cation in superconducting quantum computers,
A. Javadi-Abhari, P. Gokhale, A. Holmes, D. Franklin, K. R. Brown, M. Martonosi, and F. T. Chong, “Optimized surface code communi- cation in superconducting quantum computers,” in Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture, 2017, pp. 692–705
2017
-
[24]
A fast and high quality multilevel scheme for partitioning irregular graphs,
G. Karypis and V . Kumar, “A fast and high quality multilevel scheme for partitioning irregular graphs,” SIAM Journal on scientific Computing, vol. 20, no. 1, pp. 359–392, 1998
1998
-
[25]
Multilevelk-way partitioning scheme for irregular graphs,
——, “Multilevelk-way partitioning scheme for irregular graphs,” Jour- nal of Parallel and Distributed computing , vol. 48, no. 1, pp. 96–129, 1998
1998
-
[26]
Fluctuations of energy- relaxation times in superconducting qubits,
P. Klimov, J. Kelly, Z. Chen, M. Neeley, A. Megrant, B. Burkett, R. Barends, K. Arya, B. Chiaro, Y . Chen, A. Dunsworth, A. Fowler, B. Foxen, C. Gidney, M. Giustina, R. Graff, T. Huang, E. Jeffrey, E. Lucero, J. Mutus, O. Naaman, C. Neill, C. Quintana, P. Roushan, D. Sank, A. ...
2018
-
[27]
A game of surface codes: Large-scale quantum computing with lattice surgery,
D. Litinski, “A game of surface codes: Large-scale quantum computing with lattice surgery,” Quantum, vol. 3, p. 128, 2019
2019
-
[28]
Closing the “quantum supremacy
Y . A. Liu, X. L. Liu, F. N. Li, H. Fu, Y . Yang, J. Song, P. Zhao, Z. Wang, D. Peng, H. Chen, C. Guo, H. Huang, W. Wu, and D. Chen, “Closing the “quantum supremacy” gap: Achieving real-time simulation of a random quantum circuit using a new Sunway supercomputer,” in Proceedin...
2021
-
[29]
Simulating quantum computation by contract- ing tensor networks,
I. L. Markov and Y . Shi, “Simulating quantum computation by contract- ing tensor networks,” SIAM Journal on Computing , vol. 38, no. 3, pp. 963–981, 2008
2008
-
[30]
Quantum optimization using variational algorithms on near-term quantum devices,
N. Moll, P. Barkoutsos, L. S. Bishop, J. M. Chow, A. Cross, D. J. Egger, S. Filipp, A. Fuhrer, J. M. Gambetta, M. Ganzhorn, A. Kandala, A. Mezzacapo, P. M ¨uller, W. Riess, G. Salis, J. Smolin, I. Tavernelli, and K. Temme, “Quantum optimization using variational algorithms on ...
2018
-
[31]
Suppression of qubit crosstalk in a tunable coupling superconducting circuit,
P. Mundada, G. Zhang, T. Hazard, and A. Houck, “Suppression of qubit crosstalk in a tunable coupling superconducting circuit,”Physical Review Applied, vol. 12, no. 5, p. 054023, 2019
2019
-
[32]
Noise-adaptive compiler mappings for noisy intermediate-scale quan- tum computers,
P. Murali, J. M. Baker, A. Javadi-Abhari, F. T. Chong, and M. Martonosi, “Noise-adaptive compiler mappings for noisy intermediate-scale quan- tum computers,” in Proceedings of the Twenty-Fourth International 13 Conference on Architectural Support for Programming Languages and ...
2019
-
[33]
Software mitigation of crosstalk on noisy intermediate-scale quantum computers,
P. Murali, D. C. McKay, M. Martonosi, and A. Javadi-Abhari, “Software mitigation of crosstalk on noisy intermediate-scale quantum computers,” in Proceedings of the Twenty-Fifth International Conference on Archi- tectural Support for Programming Languages and Operating Systems ...
2020
-
[34]
Optimal qubit assignment and routing via integer programming,
G. Nannicini, L. S. Bishop, O. Gunluk, and P. Jurcevic, “Optimal qubit assignment and routing via integer programming,” arXiv preprint arXiv:2106.06446, 2021
2021 arXiv
-
[35]
A practical introduction to tensor networks: Matrix product states and projected entangled pair states,
R. Or ´us, “A practical introduction to tensor networks: Matrix product states and projected entangled pair states,” Annals of Physics , vol. 349, pp. 117–158, 2014
2014
-
[36]
Simulating large quantum circuits on a small quantum computer,
T. Peng, A. W. Harrow, M. Ozols, and X. Wu, “Simulating large quantum circuits on a small quantum computer,” Physical Review Letters , vol. 125, no. 15, p. 150504, 2020
2020
-
[37]
Circuit knitting with classical communica- tion,
C. Piveteau and D. Sutter, “Circuit knitting with classical communica- tion,” IEEE Transactions on Information Theory , 2023
2023
-
[38]
Graph minors. x. obstructions to tree- decomposition,
N. Robertson and P. D. Seymour, “Graph minors. x. obstructions to tree- decomposition,” Journal of Combinatorial Theory, Series B , vol. 52, no. 2, pp. 153–190, 1991
1991
-
[39]
Approaches to constrained quantum approximate optimization,
Z. H. Saleem, T. Tomesh, B. Tariq, and M. Suchara, “Approaches to constrained quantum approximate optimization,” arXiv preprint arXiv:2010.06660, 2020
2010 arXiv
-
[40]
The density-matrix renormalization group in the age of matrix product states,
U. Schollw ¨ock, “The density-matrix renormalization group in the age of matrix product states,” Annals of physics , vol. 326, no. 1, pp. 96–192, 2011
2011
-
[41]
Megatron-lm: Training multi-billion parameter language models using model parallelism,
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catan- zaro, “Megatron-lm: Training multi-billion parameter language models using model parallelism,” arXiv preprint arXiv:1909.08053 , 2019
1909 arXiv
-
[42]
Polynomial-time algorithms for prime factorization and discrete logarithms on a quantum computer,
P. W. Shor, “Polynomial-time algorithms for prime factorization and discrete logarithms on a quantum computer,”SIAM review, vol. 41, no. 2, pp. 303–332, 1999
1999
-
[43]
Scaling super- conducting quantum computers with chiplet architectures,
K. N. Smith, G. S. Ravi, J. M. Baker, and F. T. Chong, “Scaling super- conducting quantum computers with chiplet architectures,” in 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2022, pp. 1092–1109
2022
-
[44]
Comparing the overhead of topological and concatenated quantum error correction,
M. Suchara, A. Faruque, C.-Y . Lai, G. Paz, F. T. Chong, and J. Ku- biatowicz, “Comparing the overhead of topological and concatenated quantum error correction,” arXiv preprint arXiv:1312.2316 , 2013
2013 arXiv
-
[45]
Optimality study of existing quantum computing layout synthesis tools,
B. Tan and J. Cong, “Optimality study of existing quantum computing layout synthesis tools,” IEEE Transactions on Computers , 2020
2020
-
[46]
Alpharouter: Quantum circuit routing with reinforcement learning and tree search,
W. Tang, Y . Duan, Y . Kharkov, R. Fakoor, E. Kessler, and Y . Shi, “Alpharouter: Quantum circuit routing with reinforcement learning and tree search,” in 2024 IEEE International Conference on Quantum Computing and Engineering (QCE) , vol. 1. IEEE, 2024, pp. 930–940
2024
-
[48]
Cutqc: using small quantum computers for large quantum circuit evaluations,
W. Tang, T. Tomesh, M. Suchara, J. Larson, and M. Martonosi, “Cutqc: using small quantum computers for large quantum circuit evaluations,” in Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , 2021, p...
2021
-
[49]
Efficient tensor network simulation of ibm’s kicked ising experiment,
J. Tindall, M. Fishman, M. Stoudenmire, and D. Sels, “Efficient tensor network simulation of ibm’s kicked ising experiment,” arXiv preprint arXiv:2306.14887, 2023
2023 arXiv
-
[50]
Coreset clustering on small quantum computers,
T. Tomesh, P. Gokhale, E. R. Anschuetz, and F. T. Chong, “Coreset clustering on small quantum computers,” Electronics, vol. 10, no. 14, p. 1690, 2021
2021
-
[51]
Matrix product density operators: Simulation of finite-temperature and dissipative systems,
F. Verstraete, J. J. Garcia-Ripoll, and J. I. Cirac, “Matrix product density operators: Simulation of finite-temperature and dissipative systems,” Physical review letters, vol. 93, no. 20, p. 207204, 2004
2004
-
[52]
Efficient classical simulation of slightly entangled quantum computations,
G. Vidal, “Efficient classical simulation of slightly entangled quantum computations,” Physical review letters, vol. 91, no. 14, p. 147902, 2003
2003
-
[53]
Efficient simulation of one-dimensional quantum many-body systems,
——, “Efficient simulation of one-dimensional quantum many-body systems,” Physical review letters, vol. 93, no. 4, p. 040502, 2004
2004
-
[54]
A single quantum cannot be cloned,
W. K. Wootters and W. H. Zurek, “A single quantum cannot be cloned,” Nature, vol. 299, no. 5886, pp. 802–803, 1982
1982
-
[55]
The surface code with a twist,
T. J. Yoder and I. H. Kim, “The surface code with a twist,” Quantum, vol. 1, p. 2, 2017
2017
-
[56]
Quantum simulation with hybrid tensor networks,
X. Yuan, J. Sun, J. Liu, Q. Zhao, and Y . Zhou, “Quantum simulation with hybrid tensor networks,” Physical Review Letters , vol. 127, no. 4, p. 040501, 2021
2021
-
[57]
Highly accurate protein structure prediction with alphafold,
A. ˇZ´ıdek, A. Potapenko, A. Bridgland, C. Meyer, S. A. A. Kohl, A. J. Ballard, A. Cowie, B. Romera-Paredes, S. Nikolov, R. Jain, J. Adler, T. Back, S. Petersen, D. Reiman, E. Clancy, M. Zielinski, M. Steinegger, M. Pacholska, T. Berghammer, S. Bodenstein, D. Silver, O. Vinyal...
2021
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.