Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Adaptive Neural Quantum States: A Recurrent Neural Network Perspective

T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Adaptive RNNs, which grow their hidden dimension during training, reach the same or better variational accuracy for quantum ground states in roughly a quarter of the wall-clock time of fixed-size RNNs, and the speedup grows with system…

desk verdict Useful and honest NQS training-speedup paper, but the accuracy claims need control runs under matched learning-rate schedules before they can be trusted. read the letter →

arxiv 2507.18700 v1 pith:DNDCGEMF submitted 2025-07-24 cond-mat.dis-nn cond-mat.str-elcs.LGphysics.comp-phquant-ph

classification cond-mat.dis-nncond-mat.str-elcs.LGphysics.comp-phquant-ph
keywords neuralquantumstatesrecurrentnetworksvariationalMonteCarloadaptivetraininghidden-stategrowthparameterpaddingmany-bodygroundGPUefficiency
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that variational calculations of quantum ground states with recurrent-neural-network wave functions become much cheaper if the network's hidden memory grows during training instead of staying fixed. The proposed Adaptive scheme trains a small RNN first, then pads its weights with small random numbers to initialize a larger RNN, repeating until the target size is reached. On four spin-model benchmarks the scheme finishes in a fraction of the wall-clock time, such as 34% on the 100-spin 1D transverse-field Ising model with an asymptotic ratio near 25.6%, and about 6.5 hours versus 14.8 hours on the 6x6 2D Heisenberg model, while matching or improving variational energies and variances. The reason to care is that if true, larger and more accurate neural quantum states become accessible within existing GPU budgets.

What carries the argument

The load-bearing mechanism is parameter padding at each capacity increase: when the hidden dimension doubles, existing weight matrices and biases are copied into the larger tensors and the remaining entries are filled with small random numbers, so the larger RNN begins as a perturbed continuation of the smaller one rather than a fresh random initialization. Training is organized as a sequence of GRU-based RNN quantum states, each pretrained by the previous one; the hidden-state memory controls expressivity, the chain rule gives autoregressive sampling, and the optimizer's momentum is carried across transfers. This is what converts training a big network into training a small network and then growing it.

What would settle it

Run a Static RNN with the final hidden dimension using exactly the Adaptive learning-rate schedule, such as 5e-3 for the first half of training and 5e-4 for the second half, and compare wall-clock time, final energy, and variance; if the Static model matches the Adaptive results, the benefit would come from the schedule rather than from growing the hidden dimension.

Watch

Extended reading notes

Core claim

The central claim is that a capacity-growing training scheme makes RNN wave functions cheaper and more stable without sacrificing accuracy. Starting from a small hidden dimension, the network is trained for a fixed number of steps, then its hidden dimension is doubled and its weights and biases are padded with small random numbers to initialize the larger network; the optimizer's momentum is carried over, and all parameters remain trainable. Applied to ground-state searches via variational Monte Carlo, this Adaptive RNN reaches the same or better energies and variances than a Static RNN of the final size trained for the same number of steps, while using a fraction of the wall-clock time. On the long-range TFIM and cluster-state problems it reaches lower final energies than Static RNNs, which the authors interpret as better avoidance of local minima, and a mid-sequence model with half the final hidden dimension matches the Static model's accuracy.

Load-bearing premise

The load-bearing premise is that the gains come from the capacity-growing scheme itself and not from the different learning-rate schedules used for the Static and Adaptive runs, since in three of four benchmarks the schedules differ and no Static control with the Adaptive schedule is reported.

Editorial extensions

If this is right

  • On the 1D TFIM with 100 spins, the Adaptive RNN finishes in 34% of the Static RNN's wall-clock time and reaches comparable or lower energy variance, with the fitted time ratio approaching about 25.6% as system size grows.
  • On the 6x6 2D Heisenberg model, the Adaptive RNN reaches the best energy, error, and variance in 6 hours 33 minutes versus 14 hours 48 minutes for the Static RNN, and the early-stopping variant also beats the Static model while taking 13 hours 49 minutes.
  • On the long-range TFIM and the 1D cluster state, the Adaptive RNN reaches lower final variational energies than the Static RNN, suggesting it avoids excited-state local minima that trap fixed-size training.
  • Models earlier in the Adaptive sequence with half the final hidden dimension achieve energies comparable to the Static model with the full dimension, indicating better trainability with fewer parameters.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the benefit is a curriculum effect, where the network first learns simple low-entanglement structure before adding capacity, then the optimal growth schedule should track the entanglement structure of the target state; a testable extension would vary the doubling interval and compare across different phases.
  • The padding trick is architecture-agnostic in principle, so the same small-to-large idea could be tried on other neural quantum states such as transformers or convolutional ansatze, but the paper only demonstrates it for GRU-based RNNs, leaving transfer to other architectures an open question.
  • Because the penultimate half-size model already matches the full-size Static model's accuracy, the scheme may also serve as a parameter-reduction strategy, lowering GPU memory pressure on larger lattices even when wall-clock time is not the main constraint.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes an Adaptive training scheme for recurrent neural network (RNN) quantum states, in which the hidden-state dimension is doubled at fixed intervals and the parameters of the smaller model, along with the Adam optimizer state, are padded with small random values to initialize the larger model. The method is benchmarked against fixed-size (Static) RNNs on four spin Hamiltonians: the 1D transverse-field Ising model (TFIM), the 2D square-lattice Heisenberg model, the long-range TFIM, and the 1D cluster state. The authors report that Adaptive RNNs reach comparable or better variational energies and energy variances while training in a fraction of the wall-clock time in the 1D TFIM and 2D Heisenberg benchmarks, with an asymptotic time ratio of about 25.6% for large 1D systems. They also report improved stability of the training trajectory and argue that the scheme helps avoid local minima, particularly for the cluster state and long-range TFIM.

Significance. The speed advantage of the Adaptive scheme is well supported: the wall-clock measurements across four benchmarks consistently show that training small models early and reusing them as initializations saves substantial time, and the asymptotic scaling analysis in Appendix C is a useful quantitative contribution. The paper also ships open-source code, which is a valuable strength for reproducibility. If the accuracy and stability improvements are confirmed, the method would be a practical and broadly applicable improvement for NQS optimization, analogous to growing the bond dimension in DMRG. However, the current evidence for improved accuracy and stability is not cleanly separated from unequal learning-rate schedules between Static and Adaptive runs, and the stability claim rests on single training trajectories. These issues are load-bearing for the paper's central claims and require additional controlled experiments.

major comments (3)
  1. [§III.A, §III.B, §III.D and Table III] The accuracy and stability comparison is confounded by different learning-rate schedules. In the 1D TFIM, Static uses a fixed 5e-4 while Adaptive uses 5e-3 for the first 25,000 steps and 5e-4 afterwards; in the 2D Heisenberg model the two schedules differ (Static decay from 5e-4, Adaptive fixed 5e-4 followed by a decay); and in the cluster state Static uses 1e-4 while Adaptive uses 1e-3. No Static control run with the Adaptive learning-rate schedule is reported. Because the Adam momentum is carried across growth steps, the entire optimization trajectory differs, not just the final learning rate. A direct control—for example, a Static dh=256 run under exactly the Adaptive learning-rate schedule, or matched schedules for both methods—is required before the energy/variance improvements can be attributed to the growing hidden-dimension scheme rather than to learning-rate tuning.
  2. [§III.A and Appendix D] The claim that Adaptive training 'reduces training fluctuations' is not supported by the data presented. Figures 3(a) and 7 show one Static and one Adaptive trajectory per system size; the error bars in Tables I and II are Monte Carlo sampling errors on the final 1,000,000-sample estimates, not seed-to-seed training variability. A single trajectory cannot establish a reduction in training fluctuations. The authors should either provide multiple independent training runs with error bars on the training curves and final observables, or explicitly restrict the stability claim to the particular trajectories shown.
  3. [Appendix B, Fig. 5 caption, and Table III] There are internal inconsistencies in the reported hyperparameters that affect reproducibility. The text in §III.A and Table III state the Static 1D TFIM learning rate is 5e-4, but the Fig. 5 caption in Appendix B says '10^-4 is identified as the optimal rate'. Also, Table III lists the 2D Heisenberg Adaptive model as 'dh doubling every 25,000 steps', while the text in §III.B states 'double it every 50,000 steps'. These inconsistencies should be reconciled, since the empirical comparison depends on the precise training schedule.
minor comments (4)
  1. [Fig. 6 caption] The caption contains a typo: 'Atatic' should be 'Static'.
  2. [Appendix A] The mapping for the GLU layer writes 'b1, b2 ∈∈ R^{dmodel}'; the double '∈' is a typo and should be corrected.
  3. [§III.D] The sentence 'The latter is a good indicator of the quality of a variational calculation' uses 'latter' to refer to the energy variance; consider clarifying the antecedent in the surrounding discussion.
  4. [§III.C] The interpretation that Adaptive RNNs start with 'a low entanglement structure' is plausible but presented without direct entanglement measurements; a brief caveat or reference to the entanglement properties of small-dh RNNs would strengthen the discussion.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the paper is an empirical comparison against external benchmarks, with no equation-level reduction or fitted parameter renamed as a prediction.

full rationale

The central claims of the paper are empirical: an Adaptive RNN training schedule is compared against a Static RNN on several Hamiltonians, with runtimes, variational energies, and variances measured against externally grounded references (DMRG energies and the exact cluster-state energy of -64 from Ref. [72]). The RNN and GRU constructions are taken from prior work (Refs. [25, 55]) and used as tools, not as evidence for the new adaptive-scheme claims. The wall-clock speedup is reported as a measured quantity, and the asymptotic ratio of 25.6% in Appendix C is a fitted extrapolation of measured runtimes rather than a derived prediction that reduces to its inputs. The accuracy comparisons in three benchmarks use different learning-rate schedules for Static and Adaptive models, which is a potential experimental confound and a correctness risk, but it is not a circularity: no equation is defined in terms of its own output, and no fitted parameter is relabeled as a prediction. The explanatory remarks about low-entanglement structure cite external results and are hypotheses, not load-bearing derivations. Therefore no circular step is present under the definitions used here.

Assumptions & free parameters 7 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new physical entities; the adaptive scheme is a training procedure. The central claim rests on hand-chosen hyperparameters (learning rates, growth intervals, early-stopping thresholds) and on external references (DMRG, Ref. 72) for accuracy assessment.

free parameters (7)
  • Static 1D TFIM learning rate = 5e-4
    Selected by scanning 1e-5 to 1e-3; the baseline runs use this fixed rate.
  • Adaptive 1D TFIM learning rate schedule = 5e-3 until 25k steps, then 5e-4
    Chosen by trial over pairs; differs from the Static baseline, confounding the comparison.
  • Hidden-dimension growth interval (1D TFIM) = 6250 steps per doubling
    Chosen by hand; regular doubling schedule from dh=2 to dh=256.
  • 2D Heisenberg Static learning rate schedule = 5e-4 * (1 + t/5000)^-1
    Exponential decay schedule selected for the Static baseline.
  • 2D Heisenberg Adaptive learning rate schedule = 5e-4 * (1 + (t-100k)/5000 * floor(t/100k))^-1
    Piecewise decay schedule; differs from the Static schedule.
  • Cluster state learning rates = Static 1e-4, Adaptive 1e-3
    Best of three tested rates for each method; not matched across methods.
  • Early-stopping delta and patience (2D Heisenberg) = delta = 10^(-1/2 log2(dh)), patience = 10000
    Hand-chosen criterion for the early-stopping variant.
assumptions (4)
  • domain assumption VMC energy minimization converges to the ground state for sufficiently expressive ansatze.
    The paper relies on the variational principle throughout (Sec. III introduction).
  • domain assumption DMRG reference energies are accurate ground-state energies.
    Used to define relative errors in Tables I and II (Sec. III.A).
  • domain assumption Marshall sign rule maps the 2D Heisenberg ground state to a positive amplitude state.
    Invoked in Sec. III.B to justify using a positive RNN.
  • domain assumption The RNN/GRU parameterization is expressive enough to represent the targeted ground states.
    Assumed throughout; the adaptive scheme changes only training dynamics, not representation power.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Adaptive Neural Quantum States: A Recurrent Neural Network Perspective." pith.science (2026). https://pith.science/paper/DNDCGEMF

@misc{pith2026250718700,
  author       = {Pith},
  title        = {Pith review of: Adaptive Neural Quantum States: A Recurrent Neural Network Perspective},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DNDCGEMF}},
  note         = {Machine review of arXiv:2507.18700}
}
read the original abstract

Neural-network quantum states (NQS) are powerful neural-network ans\"atzes that have emerged as promising tools for studying quantum many-body physics through the lens of the variational principle. These architectures are known to be systematically improvable by increasing the number of parameters. Here we demonstrate an Adaptive scheme to optimize NQSs, through the example of recurrent neural networks (RNN), using a fraction of the computation cost while reducing training fluctuations and improving the quality of variational calculations targeting ground states of prototypical models in one- and two-spatial dimensions. This Adaptive technique reduces the computational cost through training small RNNs and reusing them to initialize larger RNNs. This work opens up the possibility for optimizing graphical processing unit (GPU) resources deployed in large-scale NQS simulations.

Figures

Figures reproduced from arXiv: 2507.18700 by the authors.

Figure 1
Figure 1. FIG. 1. (a) An illustration of a recurrent neural network (RNN) in the rolled version on the left-hand side and the unrolled [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2. (a) Each RNN model (green box) in the sequence is pretrained by the previous model, and inherits the parameters. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3. (a) Energy variance per spin throughout training for [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: FIG. 4. A comparison in terms of the variational energy be [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: FIG. 5. Variance throughout training of the Static framework for all system sizes of the 1D TFIM that were studied. Five [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: FIG. 6. Comparison of time taken for training the Atatic [PITH_FULL_IMAGE:figures/full_fig_p012_6.png]
Figure 7
Figure 7. Figure 7: FIG. 7. Variance per spin throughout training for the 1D TFIM benchmark across all system sizes not shown in manuscript. [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bridging Frustration and Non-Hermiticity via COMPASS: An Adaptive Biorthogonal Neural Quantum State Framework

    quant-ph 2026-07 conditional novelty 6.0 of 10

    COMPASS combines adaptive recurrent neural networks with biorthogonal Monte Carlo to simulate non-Hermitian spin systems, revealing ansatz-induced PT breaking, frustration-gap shielding, and a diabolic ring of level c...

Reference graph

Works this paper leans on

79 extracted references · 33 canonical work pages · cited by 1 Pith paper

  1. [1]

    Carleo, I

    G. Carleo, I. Cirac, K. Cranmer, L. Daudet, M. Schuld, N. Tishby, L. Vogt-Maranto, and L. Zdeborov´ a, Machine learning and the physical sciences, Rev. Mod. Phys. 91, 045002 (2019)

  2. [2]

    Dawid, J

    A. Dawid, J. Arnold, B. Requena, A. Gresch, M. Podzie, K. Donatella, K. A. Nicoli, P. Stornati, R. Koch, M. Bt- tner, R. Okua, G. Muoz-Gil, R. A. Vargas-Hernndez, A. Cervera-Lierta, J. Carrasquilla, V. Dunjko, M. Gabri, P. Huembeli, E. van Nieuwenburg, F. Vicentini, L. Wang, S. J. Wetzel, G. Carleo, E. Greplov, R. Krems, F. Mar- quardt, M. Tomza, M. Lewen...

  3. [3]

    R. G. Melko and J. Carrasquilla, Language models for quantum simulation, Nature Computational Science 4, 11 (2024)

  4. [4]

    Androsiuk, L

    J. Androsiuk, L. Kuak, and K. Sienicki, Neural network solution of the schrdinger equation for a two-dimensional harmonic oscillator, Chemical Physics 173, 377 (1993)

  5. [5]

    Lagaris, A

    I. Lagaris, A. Likas, and D. Fotiadis, Artificial neural net- work methods in quantum mechanics, Computer Physics Communications 104, 1 (1997)

  6. [6]

    Sugawara, Numerical solution of the schrdinger equa- tion by neural network and genetic algorithm, Computer Physics Communications 140, 366 (2001)

    M. Sugawara, Numerical solution of the schrdinger equa- tion by neural network and genetic algorithm, Computer Physics Communications 140, 366 (2001)

  7. [7]

    Carleo and M

    G. Carleo and M. Troyer, Solving the quantum many- body problem with artificial neural networks, Science 355, 602 (2017), publisher: American Association for the Advancement of Science

  8. [8]

    Lange, A

    H. Lange, A. Van de Walle, A. Abedinnia, and A. Bohrdt, From architectures to applications: a review of neural quantum states, Quantum Science and Technology 9, 040501 (2024)

Show all 79 references
  1. [9]

    Medvidovi´ c and J

    M. Medvidovi´ c and J. R. Moreno, Neural-network quan- tum states for many-body physics, The European Phys- ical Journal Plus 139, 631 (2024)

  2. [10]

    Nomura and M

    Y. Nomura and M. Imada, Dirac-type nodal spin liq- uid revealed by refined quantum many-body solver using neural-network wave function, correlation ratio, and level spectroscopy, Phys. Rev. X 11, 031034 (2021)

  3. [11]

    Hibat-Allah, R

    M. Hibat-Allah, R. G. Melko, and J. Carrasquilla, Sup- plementing recurrent neural network wave functions with symmetry and annealing to improve accuracy (2024), arXiv:2207.14314 [cond-mat.dis-nn]

  4. [12]

    Chen and M

    A. Chen and M. Heyl, Empowering deep neural quantum states through efficient optimization, Nature Physics 20, 1476 (2024)

  5. [13]

    Rende, L

    R. Rende, L. L. Viteritti, L. Bardone, F. Becca, and S. Goldt, A simple linear algebra identity to optimize large-scale neural network quantum states, Communica- tions Physics 7, 10.1038/s42005-024-01732-4 (2024)

  6. [14]

    W. M. C. Foulkes, L. Mitas, R. J. Needs, and G. Ra- jagopal, Quantum monte carlo simulations of solids, Rev. Mod. Phys. 73, 33 (2001)

  7. [15]

    S. R. White, Density matrix formulation for quantum renormalization groups, Phys. Rev. Lett.69, 2863 (1992)

  8. [16]

    Verstraete, T

    F. Verstraete, T. Nishino, U. Schollw¨ ock, M. C. Ba˜ nuls, G. K. Chan, and M. E. Stoudenmire, Density matrix renormalization group, 30 years on, Nature Reviews Physics 5, 273 (2023)

  9. [17]

    Mareschal, The early years of quantum monte carlo (1): the ground state, The European Physical Journal H 46, 10.1140/epjh/s13129-021-00017-6 (2021)

    M. Mareschal, The early years of quantum monte carlo (1): the ground state, The European Physical Journal H 46, 10.1140/epjh/s13129-021-00017-6 (2021)

  10. [18]

    W. L. McMillan, Ground state of liquid he 4, Phys. Rev. 138, A442 (1965)

  11. [19]

    Becca and S

    F. Becca and S. Sorella, Quantum Monte Carlo Ap- proaches for Correlated Systems (Cambridge University 12 10 6 10 5 10 4 10 3 10 2 10 1 100 σ2/N N = 20 N = 40 N = 60 0 10000 20000 30000 40000 50000 Training Step 10 6 10 5 10 4 10 3 10 2 10 1 100 σ2/N N = 80 0 10000 20000 30000...

  12. [20]

    Schmitt and M

    M. Schmitt and M. Heyl, Quantum many-body dynamics in two dimensions with artificial neural networks, Physi- cal Review Letters 125, 10.1103/physrevlett.125.100503 (2020)

  13. [21]

    Nomura, A

    Y. Nomura, A. S. Darmawan, Y. Yamaji, and M. Imada, Restricted boltzmann machine learning for solving strongly correlated quantum systems, Phys. Rev. B 96, 205152 (2017)

  14. [22]

    Cai and J

    Z. Cai and J. Liu, Approximating quantum many-body wave functions using artificial neural networks, Phys. Rev. B 97, 035116 (2018)

  15. [23]

    K. Choo, G. Carleo, N. Regnault, and T. Neupert, Sym- metries and many-body excitations with neural-network quantum states, Phys. Rev. Lett. 121, 167204 (2018)

  16. [24]

    K. Choo, T. Neupert, and G. Carleo, Two-dimensional frustrated J1−J2 model studied with neural network quantum states, Phys. Rev. B 100, 125124 (2019)

  17. [25]

    Hibat-Allah, M

    M. Hibat-Allah, M. Ganahl, L. E. Hayward, R. G. Melko, and J. Carrasquilla, Recurrent neural network wave func- tions, Physical Review Research 2, 10.1103/physrevre- search.2.023358 (2020)

  18. [26]

    Roth, Iterative Retraining of Quantum Spin Models Using Recurrent Neural Networks (2020), arXiv:2003.06228

    C. Roth, Iterative Retraining of Quantum Spin Models Using Recurrent Neural Networks (2020), arXiv:2003.06228

  19. [27]

    Casert, T

    C. Casert, T. Vieijra, S. Whitelam, and I. Tamblyn, Dy- namical large deviations of two-dimensional kinetically constrained models using a neural-network state ansatz, Phys. Rev. Lett. 127, 120602 (2021)

  20. [28]

    D. Luo, Z. Chen, K. Hu, Z. Zhao, V. M. Hur, and B. K. Clark, Gauge-invariant and anyonic-symmetric au- toregressive neural network for quantum lattice models, Phys. Rev. Res. 5, 013216 (2023)

  21. [30]

    M. S. Moss, R. Wiersema, M. Hibat-Allah, J. Car- rasquilla, and R. G. Melko, Leveraging recurrence in neu- ral network wavefunctions for large-scale simulations of heisenberg antiferromagnets: the square lattice (2025), arXiv:2502.17144 [cond-mat.str-el]. 13 0 10000 20000 3000...

  22. [31]

    M. S. Moss, R. Wiersema, M. Hibat-Allah, J. Car- rasquilla, and R. G. Melko, Leveraging recurrence in neural network wavefunctions for large-scale simulations of heisenberg antiferromagnets: the triangular lattice (2025), arXiv:2505.20406 [cond-mat.str-el]

  23. [33]

    Sprague and S

    K. Sprague and S. Czischek, Variational monte carlo with large patched transformers, Communications Physics 7, 10.1038/s42005-024-01584-y (2024)

  24. [34]

    J. A. Sobral, M. Perle, and M. S. Scheurer, Physics- informed transformers for electronic quantum states (2024), arXiv:2412.12248 [cond-mat.str-el]

  25. [35]

    Lange, G

    H. Lange, G. Bornet, G. Emperauger, C. Chen, T. La- haye, S. Kienle, A. Browaeys, and A. Bohrdt, Trans- former neural networks and quantum simulators: a hy- brid approach for simulating strongly correlated systems, Quantum 9, 1675 (2025)

  26. [36]

    X. Hu, L. Chu, J. Pei, W. Liu, and J. Bian, Model com- plexity of deep learning: a survey, Knowledge and Infor- mation Systems 63, 25852619 (2021)

  27. [37]

    Schollwck, The density-matrix renormalization group in the age of matrix product states, Annals of Physics 326, 96192 (2011)

    U. Schollwck, The density-matrix renormalization group in the age of matrix product states, Annals of Physics 326, 96192 (2011)

  28. [38]

    Ganahl, J

    M. Ganahl, J. Beall, M. Hauru, A. G. Lewis, T. Wo- jno, J. H. Yoo, Y. Zou, and G. Vidal, Density matrix renormalization group with tensor processing units, PRX Quantum 4, 010317 (2023)

  29. [39]

    S. J. Pan and Q. Yang, A survey on transfer learning, IEEE Transactions on Knowledge and Data Engineering 22, 1345 (2010)

  30. [40]

    A. A. Rusu, N. C. Rabinowitz, G. Desjardins, H. Soyer, J. Kirkpatrick, K. Kavukcuoglu, R. Pascanu, and R. Hadsell, Progressive neural networks (2022), arXiv:1606.04671 [cs.LG]

  31. [41]

    E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, Lora: Low-rank adap- tation of large language models (2021), arXiv:2106.09685 [cs.CL]

  32. [42]

    T. Chen, I. Goodfellow, and J. Shlens, Net2net: Accelerating learning via knowledge transfer (2016), arXiv:1511.05641 [cs.LG]

  33. [43]

    R. Zen, L. My, R. Tan, F. Hbert, M. Gattobigio, C. Miniatura, D. Poletti, and S. Bressan, Transfer learn- ing for scalability of neural-network quantum states, Physical Review E 101, 10.1103/physreve.101.053301 (2020)

  34. [44]

    Zen and S

    R. Zen and S. Bressan, Transfer learning for larger, broader, and deeper neural-network quantum states, in Database and Expert Systems Applications , edited by C. Strauss, G. Kotsis, A. M. Tjoa, and I. Khalil (Springer International Publishing, Cham, 2021) pp. 207–219

  35. [45]

    D. Wu, R. Rossi, F. Vicentini, and G. Carleo, From tensor-network quantum states to tensorial re- current neural networks, Physical Review Research 5, 10.1103/physrevresearch.5.l032001 (2023)

  36. [46]

    Kelley, Sequence modeling with recurrent tensor net- works (2016)

    R. Kelley, Sequence modeling with recurrent tensor net- works (2016)

  37. [47]

    Hibat-Allah, E

    M. Hibat-Allah, E. M. Inack, R. Wiersema, R. G. Melko, and J. Carrasquilla, Variational neural annealing, Nature Machine Intelligence 3, 952961 (2021)

  38. [48]

    S. R. White, Density-matrix algorithms for quantum renormalization groups, Phys. Rev. B 48, 10345 (1993)

  39. [49]

    Rder, and B

    Legeza, J. Rder, and B. A. Hess, Controlling the accuracy of the density-matrix renormalization-group method: The dynamical block state selection approach, Physical Review B 67, 10.1103/physrevb.67.125114 (2003)

  40. [50]

    Z. C. Lipton, J. Berkowitz, and C. Elkan, A critical re- view of recurrent neural networks for sequence learning (2015)

  41. [51]

    A. M. Sch¨ afer and H. G. Zimmermann, Recurrent neural networks are universal approximators, in Artificial Neu- ral Networks – ICANN 2006 , edited by S. D. Kollias, A. Stafylopatis, W. Duch, and E. Oja (Springer Berlin Heidelberg, Berlin, Heidelberg, 2006) pp. 632–640

  42. [52]

    G. S. Carmantini, P. b. Graben, M. Desroches, and S. Ro- drigues, Turing computation with recurrent artificial neu- ral networks (2015)

  43. [53]

    Carrasquilla, G

    J. Carrasquilla, G. Torlai, R. G. Melko, and L. Aolita, Reconstructing quantum states with generative models, Nature Machine Intelligence 1, 155 (2019)

  44. [54]

    Hibat-Allah, R

    M. Hibat-Allah, R. G. Melko, and J. Carrasquilla, In- vestigating topological order using recurrent neural net- works, Phys. Rev. B 108, 075152 (2023)

  45. [55]

    K. Cho, B. van Merrienboer, C. Gulcehre, D. Bah- danau, F. Bougares, H. Schwenk, and Y. Bengio, Learn- ing phrase representations using rnn encoder-decoder for statistical machine translation (2014), arXiv:1406.1078 [cs.CL]

  46. [56]

    Bravyi, D

    S. Bravyi, D. P. Divincenzo, R. Oliveira, and B. M. Ter- hal, The complexity of stoquastic local hamiltonian prob- 14 lems, Quantum Info. Comput. 8, 361 (2008)

  47. [57]

    D. P. Kingma and J. Ba, Adam: A method for stochastic optimization (2014), arXiv:1412.6980 [cs.LG]

  48. [58]

    B. A. Cipra, An introduction to the ising model, The American Mathematical Monthly 94, 937 (1987)

  49. [59]

    Assaraf and M

    R. Assaraf and M. Caffarel, Zero-variance zero-bias prin- ciple for observables in quantum monte carlo: Applica- tion to forces, The Journal of Chemical Physics 119, 1053610552 (2003)

  50. [60]

    D. Wu, R. Rossi, F. Vicentini, N. Astrakhantsev, F. Becca, X. Cao, J. Carrasquilla, F. Ferrari, A. Georges, M. Hibat-Allah, M. Imada, A. M. Luchli, G. Mazzola, A. Mezzacapo, A. Millis, J. R. Moreno, T. Neupert, Y. Nomura, J. Nys, O. Parcollet, R. Pohle, I. Romero, M. Schmid, J...

  51. [61]

    A. M. Mood, Introduction to the Theory of Statistics (Theorem 2). (McGraw-hill, 1950)

  52. [62]

    A. W. Sandvik and J. Kurkijrvi, Quantum Monte Carlo simulation method for spin systems, Physical Review B 43, 5950 (1991)

  53. [63]

    Liu and E

    Z. Liu and E. Manousakis, Variational calculations for the square-lattice quantum antiferromagnet, Physical Review B 40, 11437 (1989)

  54. [64]

    S. R. White and A. L. Chernyshev, Nel Order in Square and Triangular Lattice Heisenberg Models, Physical Re- view Letters 99, 127004 (2007)

  55. [65]

    Marshall, Antiferromagnetism, Proceedings of the Royal Society of London

    W. Marshall, Antiferromagnetism, Proceedings of the Royal Society of London. Series A, Mathematical and Physical Sciences 232, 48 (1955), publisher: The Royal Society

  56. [66]

    Capriotti, Quantum Effects and Broken Symme- tries in Frustrated Antiferromagnets (2001), arXiv:cond- mat/0112207

    L. Capriotti, Quantum Effects and Broken Symme- tries in Frustrated Antiferromagnets (2001), arXiv:cond- mat/0112207

  57. [67]

    Zhang, G

    J. Zhang, G. Pagano, P. W. Hess, A. Kyprianidis, P. Becker, H. Kaplan, A. V. Gorshkov, Z.-X. Gong, and C. Monroe, Observation of a many-body dynamical phase transition with a 53-qubit quantum simulator, Nature 551, 601604 (2017)

  58. [68]

    Bernien, S

    H. Bernien, S. Schwartz, A. Keesling, H. Levine, A. Om- ran, H. Pichler, S. Choi, A. S. Zibrov, M. Endres, M. Greiner, V. Vuleti, and M. D. Lukin, Probing many- body dynamics on a 51-atom quantum simulator, Nature 551, 579584 (2017)

  59. [69]

    Koffel, M

    T. Koffel, M. Lewenstein, and L. Tagliacozzo, Entangle- ment entropy for the long-range ising chain in a trans- verse field, Physical Review Letters 109, 10.1103/phys- revlett.109.267203 (2012)

  60. [70]

    Defenu, T

    N. Defenu, T. Donner, T. Macr, G. Pagano, S. Ruffo, and A. Trombettoni, Long-range interacting quantum sys- tems, Reviews of Modern Physics 95, 10.1103/revmod- phys.95.035002 (2023)

  61. [71]

    Levine, O

    Y. Levine, O. Sharir, N. Cohen, and A. Shashua, Quan- tum entanglement in deep learning architectures, Phys. Rev. Lett. 122, 065301 (2019)

  62. [72]

    T.-H. Yang, M. Soleimanifar, T. Bergamaschi, and J. Preskill, When can classical neural networks represent quantum states? (2024), arXiv:2410.23152 [quant-ph]

  63. [73]

    Raussendorf and H

    R. Raussendorf and H. Briegel, Computational model underlying the one-way quantum computer (2002), arXiv:quant-ph/0108067 [quant-ph]

  64. [74]

    A. C. Doherty and S. D. Bartlett, Identifying phases of quantum many-body systems that are universal for quantum computation, Physical Review Letters 103, 10.1103/physrevlett.103.020506 (2009)

  65. [75]

    Bukov, M

    M. Bukov, M. Schmitt, and M. Dupont, Learning the ground state of a non-stoquastic quantum hamiltonian in a rugged neural network landscape, SciPost Physics 10, 10.21468/scipostphys.10.6.147 (2021)

  66. [76]

    Nomura, Helping restricted boltzmann machines with quantum-state representation by restoring symmetry, Journal of Physics: Condensed Matter 33, 174003 (2021)

    Y. Nomura, Helping restricted boltzmann machines with quantum-state representation by restoring symmetry, Journal of Physics: Condensed Matter 33, 174003 (2021)

  67. [77]

    G.-B. Zhou, J. Wu, C.-L. Zhang, and Z.-H. Zhou, Min- imal gated unit for recurrent neural networks, Interna- tional Journal of Automation and Computing 13, 226 (2016)

  68. [78]

    Shen, Mutual information scaling and expressive power of sequence models (2019), arXiv:1905.04271 [cs.LG]

    H. Shen, Mutual information scaling and expressive power of sequence models (2019), arXiv:1905.04271 [cs.LG]

  69. [79]

    Hibat-Allah, E

    M. Hibat-Allah, E. Merali, G. Torlai, R. G. Melko, and J. Carrasquilla, Recurrent neural network wave func- tions for rydberg atom arrays on kagome lattice (2024), arXiv:2405.20384 [cond-mat.quant-gas]

  70. [80]

    Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier, Language modeling with gated convolutional networks (2017), arXiv:1612.08083 [cs.CL]

  71. [81]

    Shazeer, Glu variants improve transformer (2020), arXiv:2002.05202 [cs.LG]

    N. Shazeer, Glu variants improve transformer (2020), arXiv:2002.05202 [cs.LG]

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.