REVIEW 2 major objections 5 minor 67 references
Spectral Born machines learn integer data with a Fourier bias that may block overfitting even when heavily over-parameterized.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-10 23:36 UTC pith:OZVA5CLF
load-bearing objection Clean group-Fourier lift of IQP Born machines with solid classical estimators, real software, and a million-parameter RNA run; the overfitting-immunity claim is overstated relative to the evidence. the 2 major comments →
Spectral Born machines: classically trainable quantum generative models for discrete data
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Spectral Born machines—Fourier-phase unitaries on qudits—can be trained at scale on classical hardware via batched Heisenberg-Weyl moment matching under a graph-spectral MMD, and their built-in low-order spectral bias appears to prevent overfitting even when the model has orders of magnitude more parameters than training samples.
What carries the argument
Fourier-phase unitaries (a diagonal phase layer D(θ) sandwiched between quantum Fourier transforms over Z_d^n) together with classical Monte-Carlo estimation of their Heisenberg-Weyl moments, which turn the graph-spectral MMD into an efficiently optimizable loss.
Load-bearing premise
That sampling the trained model stays classically hard once parameters are small and the distribution is spectrally concentrated—the regime the paper itself flags as open to possible dequantization.
What would settle it
Classically sample (or closely approximate) the output distribution of a trained spectral Born machine whose low-order Fourier coefficients match real data, using only the polynomial number of moments accessible during classical training; success would remove the quantum-deploy advantage.
If this is right
- Integer-structured generative tasks can be attacked with far smaller parameter counts by restricting the phase gates to low-degree Fourier generators.
- Million-parameter quantum generative models become trainable today on classical GPUs, with quantum hardware needed only at inference.
- Over-parameterization can be used aggressively for expressivity without the usual overfitting penalty when the model class is spectrally biased.
- Graph-spectral MMD losses give a principled way to encode whether discrete variables are ordinal or purely categorical.
- The same classical training loop can be scaled to distributed clusters, opening billion-parameter quantum generative models.
Where Pith is reading between the lines
- The same Fourier-phase construction may supply a controllable spectral regularizer for other discrete generative families beyond Born machines.
- If dequantization via low-order Fourier sparsity fails, spectral Born machines occupy a practically useful complexity middle ground where hardness is neither proved nor easily spoofed.
- Matching the gate-set degrees to the cyclic distance of the target may become a standard inductive-bias design rule for any Fourier-based quantum generative model.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces spectral Born machines as a generalization of IQP Born machines to qudits via the group Fourier transform over Z_d^n. A Fourier-phase unitary (QFT layer, diagonal parameterized phase generated by Heisenberg-Weyl operators, inverse QFT) defines a generative model over integer vectors whose computational-basis distribution can be trained classically: Proposition 1 and the batched linear-algebra estimator (Apps. A–B) show that HW expectation values are Monte-Carlo estimable to inverse-polynomial additive error, enabling an unbiased graph-spectral MMD loss (App. C) based on cycle or complete graphs. The construction is released in the PennyLane tcdq module. Numerical experiments on a Potts model demonstrate that restricting gate degree yields comparable test MMD with far fewer parameters than the unrestricted gate set; a 190-qubit (95-qudit) model with >1 M parameters is trained on a scarce 93-nt rRNA data set and matches low-order Fourier coefficients to within the 1/√|X| floor. The authors interpret the latter as evidence that extreme over-parameterization may be immune to overfitting because of the built-in low-order spectral bias.
Significance. If the classical-trainability pipeline and the spectral inductive bias hold, the work supplies a practical route to large-scale quantum generative modeling for non-binary discrete data, removing the need for quantum gradient estimation and data loading during training. The clean Monte-Carlo theory (Prop. 1, Apps. A–C), the explicit batched estimator, and the public tcdq implementation are concrete engineering contributions that enable empirical exploration at the 100–1000 qubit scale. The Potts results give direct evidence that matching the model’s Fourier support to the data’s smoothness structure reduces parameter count without loss of performance. The rRNA experiment, while limited, is among the largest classically trained quantum generative models reported. These elements make the manuscript a useful step toward “train classical, deploy quantum” generative models; the ultimate quantum advantage remains conjectural pending hardness results in the trained, spectrally concentrated regime.
major comments (2)
- [abstract; Sec. V B; Fig. 5; Sec. VI A] The central interpretive claim that “highly over-parameterized spectral Born machines may be immune to overfitting, even in strongly data-scarce regimes” (abstract; Sec. VI A) is supported only by the rRNA experiment (Sec. V B). There a ~1.3 M-parameter model reaches MMD^{2} ≈ 0.0035 (comparable to the 1/√|X| ≈ 0.05 statistical floor) and its weight-1/2 Fourier coefficients visually match the empirical spectrum (Fig. 5). Because the training loss (Eq. 42) and the complete-graph heat kernel deliberately target only low-order Heisenberg-Weyl moments, agreement on those moments does not rule out (i) memorization of the 428 training sequences via higher-order correlations invisible to the loss, or (ii) collapse onto an incorrect but spectrally smooth distribution. No generative samples, high-order moment diagnostics, or likelihood proxies are reported. The claim should be substantially quali
- [abstract; Sec. IV; Sec. VI C] The abstract asserts that the models “remain classically hard to sample from in general.” Sec. IV (paragraph following the definition of spectral Born machines) and Sec. VI C correctly note that anticoncentration/random-circuit arguments do not apply once parameters are small and the distribution is spectrally concentrated, and that dequantization via low-order Fourier sparsity is an open risk. For the practically relevant trained models that are the point of the work, classical hardness is therefore conjectural. The unqualified phrasing in the abstract and introduction should be aligned with the more cautious discussion in Sec. VI C.
minor comments (5)
- [abstract; p. 1] Abstract and opening paragraph contain missing spaces (“We presentspectral”, “newtcdqmodule”).
- [Fig. 2; Fig. 4] Fig. 2 caption claims “similar training loss, suggesting o similar test performance”; the actual test curves appear only in Fig. 4. Cross-reference or merge the figures for clarity.
- [Sec. IV B; Sec. V] The heat-kernel bandwidth t (or equivalent mean operator weight) is a free hyper-parameter; a short sensitivity study or default-selection rule would help reproducibility.
- [Sec. V B] In the rRNA experiment the three-body gates are chosen by ranking empirical Fourier coefficients of the training set. This data-dependent gate selection should be stated more prominently as part of the model-construction pipeline.
- [throughout] Typographical inconsistencies: “Z n d” vs. “Z_d^n”, occasional missing punctuation after displayed equations, and “tcdq” sometimes rendered without backticks or italics.
Circularity Check
No load-bearing circularity; classical estimability and spectral MMD follow from QFT conjugation and graph Laplacians, with only minor self-citation of the prior TCDQ/IQP training paradigm.
specific steps
-
self citation load bearing
[Sec. I (TCDQ paragraph) and Sec. IV (definition of spectral Born machines)]
"The possibility of training quantum generative models at scale in this way was first shown for instantaneous quantum polynomial (IQP) Born machines [12] … For the case d=2, spectral Born machines become the class of IQP Born machines and thus inherit similar complexity-theoretic classical hardness guarantees related to sampling [12, 43-45]."
Reference [12] shares an author (Bowles) and supplies the classical-training paradigm that is reused. The citation is not load-bearing for the new Z_d^n mathematics (QFT conjugation, HW batch estimator, graph-spectral MMD), which is derived independently; it is therefore only a minor self-citation of the surrounding framework rather than a circular reduction of the paper's central claims.
full rationale
The paper's core technical claims are self-contained derivations. Proposition 1 and Appendix A obtain the Monte-Carlo estimator of Heisenberg-Weyl expectations directly from the Fourier-phase unitary form U=(F†)D( heta)F together with the conjugation FD(k,m)F†=D(m,-k) and unit-modulus averaging; nothing is fitted or defined in terms of the target loss. The graph-spectral MMD (Eqs. 33-42) is constructed from the Laplacian of the chosen product graph (cycle or complete) and reduces to a weighted mixture of |<D(k,0)>|^{2} differences by the standard spectral theorem for circulant matrices; the bandwidth t and gate-set restrictions are free design choices, not circular redefinitions of the data. Spectral bias (Sec. IV A, Eqs. 28-29) is an immediate consequence of low-weight generators producing low-frequency phase functions. The rRNA and Potts experiments report empirical MMD values and Fourier visualizations; they do not fit a parameter on one quantity and then re-label a related quantity as a prediction. Self-citations to the authors' prior IQP/TCDQ work supply the overall train-classical-deploy-quantum template but are not used to force uniqueness or to smuggle an ansatz that itself lacks independent support. Concurrent work [30] is acknowledged as complementary. Consequently the derivation chain does not reduce any claimed result to its own inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (5)
- heat-kernel bandwidth t (or equivalent mean operator weight)
- gate-set degree / weight cut-offs and selected three-body gates
- initialization standard deviations for multi-qudit gates
- Monte-Carlo batch sizes |K| and |Z|
- learning rate and iteration count
axioms (4)
- domain assumption Sampling from IQP circuits (and by extension spectral Born machines at d=2) is classically hard under standard complexity assumptions
- domain assumption Phase function Phi_theta(z) is classically computable in poly(n,d) time for the chosen gate sets
- domain assumption Graph Laplacian heat kernel on cycle or complete graphs supplies a suitable notion of closeness for ordinal/categorical discrete data
- ad hoc to paper Low-order Fourier concentration is a desirable inductive bias that forbids memorization of sparse empirical distributions
invented entities (2)
-
spectral Born machine (Fourier-phase unitary generative model over Z_d^n)
no independent evidence
-
tcdq PennyLane module
independent evidence
read the original abstract
We present \emph{spectral Born machines}, a class of quantum generative models that results from viewing and generalizing the class of IQP Born machines through the lens of group Fourier analysis. These quantum models exploit the quantum Fourier transform to create an inductive bias that make them naturally suited to learning integer-structured data, while remaining classically hard to sample from in general. Similar to IQP Born machines, spectral Born machines can be trained efficiently at scale on classical hardware via a maximum mean discrepancy loss based on graph spectral analysis, which we make available in a new \emph{tcdq} module of the PennyLane software platform. In numerical experiments, we show how the spectral bias of the model leads to significantly reduced parameter counts compared to unstructured approaches, and demonstrate the scalability of the software by training a 190-qubit model with over 1 million parameters to successfully learn a distribution of 93 nucleotide-long ribosomal RNA. Our results suggest that highly over-parameterized spectral Born machines may be immune to overfitting, even in strongly data-scarce regimes.
Figures
Reference graph
Works this paper leans on
-
[1]
of the complex-valued functionf:G→Cand its inverse are given by ˆf(k) := X x∈G f(x)χ k(x), f(x) = 1 |G| X k∈G ˆf(k)χ ∗ k(x),(1) where the group charactersχ k :G→Ccorrespond to the one-dimensional irreducible representations ofG. We work with the product cyclic groupG=Z n d, which nat- urally relates to vectors of integers, for which the char- acters are g...
-
[2]
can be realised via a tensor product of quantum Fourier transforms ofZ d. Concretely, the QFT overZ d and its inverse are F|x⟩= 1√ d X k ωkx |k⟩,(2) F † |x⟩= 1√ d X k ω−kx |k⟩,(3) so that the quantum Fourier transform overZ n d and its inverse are F ⊗n |x⟩= 1√ dn X k∈Zn d ωk·x |k⟩,(4) (F †)⊗n |x⟩= 1√ dn X k∈Zn d ω−k·x |k⟩.(5) Note that ford= 2, we haveF=H...
-
[3]
Shift-invariant graphs: cycleC d and completeK d When eachG i is invariant under cyclic shifts of labels, which is the case for the cycle graph or the complete graph, the LaplacianL i is a circulant matrix and the eigenvectorsu k of the Laplacian are the characters of Zn d, e⊤ x uk = 1√ dn ωk·x.(39) Since the kernelK=f(L) is also diagonalised by this basi...
-
[4]
Finite-sample estimation Here we show how to efficiently estimate and train the MMD2 loss. Definep X to be an empirical distribution of the datasetXsampled from the ground truth distribu- tionp pX (x) = 1 |X | X y∈X I[y=x],(46) withIthe indicator function. Note that⟨D(k,0)⟩ pX can be computed exactly and efficiently sincep X has sparse support. Let \⟨D(k,...
-
[5]
Scaling Laws for Neural Language Models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, Scaling laws for neural language models, arXiv preprint arXiv:2001.08361 (2020)
work page internal anchor Pith review Pith/arXiv arXiv 2001
-
[6]
A. G. Wilson, Deep learning is not so mysterious or dif- ferent, arXiv preprint arXiv:2503.02113 (2025)
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[7]
A. G. Wilson and P. Izmailov, Bayesian deep learning and a probabilistic perspective of generalization, Advances in neural information processing systems33, 4697 (2020)
work page 2020
-
[8]
C. Liu, L. Zhu, and M. Belkin, Loss landscapes and op- timization in over-parameterized non-linear systems and neural networks, Applied and Computational Harmonic Analysis59, 85 (2022)
work page 2022
-
[9]
T. Hoefler, T. H¨ aner, and M. Troyer, Disentangling hype from practicality: On realistically achieving quantum ad- vantage, Communications of the ACM66, 82 (2023)
work page 2023
- [10]
- [11]
- [12]
-
[13]
K. Chinzei, S. Yamano, Q. H. Tran, Y. Endo, and H. Os- hima, Trade-off between gradient measurement efficiency and expressivity in deep quantum neural networks, npj Quantum Information11, 79 (2025)
work page 2025
-
[14]
Scalable On-Hardware Training of Quantum Neural Networks and Application to Clinical Data Imputation
N. Mathur, P. K. Barkoutsos, M. Yamada, M. Roetteler, and I. Kerenidis, Scalable on-hardware training of quan- tum neural networks and application to clinical data im- putation, arXiv preprint arXiv:2606.03517 (2026)
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[15]
Adaptive directional gradients for parameterised quantum circuits
B. Coyle, S. Raj, V. Umathe, E. A. Cher- rat, and E. Kashefi, Adaptive directional gradients for parameterised quantum circuits, arXiv preprint arXiv:2606.09734 (2026)
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[16]
E. Recio-Armengol, S. Ahmed, and J. Bowles, Train on classical, deploy on quantum: scaling generative quantum machine learning to a thousand qubits, arXiv preprint arXiv:2503.02934 (2025)
-
[17]
S. Kasture, O. Kyriienko, and V. E. Elfving, Protocols for classically training quantum generative models on probability distributions, Physical Review A108, 042406 (2023)
work page 2023
- [18]
-
[19]
Efficient training of photonic quantum generative models
F. Gottlieb, R. Mezher, B. Ventura, S. Mansfield, and A. Salavrakos, Efficient training of photonic quan- tum generative models, arXiv preprint arXiv:2603.08793 (2026)
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[20]
Z. Kolarovszki, B. Bak´ o, M. Oszmaniec, C. Oh, and Z. Zimbor´ as, Generative modeling with gaussian boson sampling: classically trainable bosonic born machines, arXiv preprint arXiv:2603.11195 (2026)
- [21]
-
[22]
Quantum Fourier Generative Models Trainable at Large Scale
C. T¨ uys¨ uz, O. Kyriienko, and M. Grossi, Quantum fourier generative models trainable at large scale, arXiv preprint arXiv:2606.28483 (2026)
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[23]
M. Herrero-Gonzalez, B. Coyle, K. McDowall, R. Grassie, S. Beentjes, A. Khamseh, and E. Kashefi, The born ul- timatum: Conditions for classical surrogation of quan- 14 tum generative models with correlators, arXiv preprint arXiv:2511.01845 (2025)
-
[24]
O. Ball´ o-Gimbernat, M. Arroyo-S´ anchez, P. Garc´ ıa- Molina, A. Garriga, and F. Vilari˜ no, Shallow instanta- neous quantum polynomial-time circuits for generative modeling on noisy intermediate-scale quantum hardware, Physical Review A113, 042617 (2026)
work page 2026
- [25]
-
[26]
J. Slim, S. Monaco, F. Rehm, D. Kr¨ ucker, and K. Borras, An iqp born machine for calorimeter image generation at 64 qubits with compiled-iqp deployment, arXiv preprint arXiv:2605.27735 (2026)
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[27]
Parity Supervision as a Driver of Generalization in Quantum Generative Modeling
M. Baumann, D. Hein, S. Udluft, T. Rohe, C. Linnhoff- Popien, and J. Stein, Parity supervision as a driver of generalization in quantum generative modeling, arXiv preprint arXiv:2605.10258 (2026)
work page internal anchor Pith review Pith/arXiv arXiv 2026
- [28]
-
[29]
Spectral methods: crucial for machine learning, natural for quantum computers?
V. Belis, J. Bowles, R. Gupta, E. Peters, and M. Schuld, Spectral methods: crucial for machine learn- ing, natural for quantum computers?, arXiv preprint arXiv:2603.24654 (2026)
work page internal anchor Pith review Pith/arXiv arXiv 2026
- [30]
-
[31]
The tcdq module is currently under development but available as part of PennyLane labs. Seehttps://docs. pennylane.ai/en/latest/code/qp_labs.html
-
[32]
PennyLane: Automatic differentiation of hybrid quantum-classical computations
V. Bergholm, J. Izaac, M. Schuld, C. Gogolin, S. Ahmed, V. Ajith, M. S. Alam, G. Alonso-Linaje, B. Akash- Narayanan, A. Asadi,et al., Pennylane: Automatic dif- ferentiation of hybrid quantum-classical computations, arXiv preprint arXiv:1811.04968 (2018)
work page internal anchor Pith review Pith/arXiv arXiv 2018
-
[33]
J. Bradbury, R. Frostig, P. Hawkins, M. J. John- son, Y. Katariya, C. Leary, D. Maclaurin, G. Nec- ula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang, JAX: composable transformations of Python+NumPy programs (2018)
work page 2018
-
[34]
R. J. Banks, A. Crippa, M. Traube, J. Unger, C. Ertler, and W. Lechner, Qudit extension of parameterized iqp circuits: A generative quantum machine learning ap- proach to integer data, arXiv preprint arXiv:2606.28236 (2026)
work page internal anchor Pith review Pith/arXiv arXiv 2026
- [35]
-
[36]
M. Krebsbach, F. Reiter, T. Wellens, H.-H. Kowal- ski, and A. Abedi, Encoding numerical data for generative quantum machine learning, arXiv preprint arXiv:2603.23407 (2026)
-
[37]
I. R. Kondor,Group theoretical methods in machine learning(Columbia University, 2008)
work page 2008
-
[38]
A. M. Childs and W. Van Dam, Quantum algorithms for algebraic problems, Reviews of Modern Physics82, 1 (2010)
work page 2010
-
[39]
A. Asadian, P. Erker, M. Huber, and C. Kl¨ ockl, Heisenberg-weyl observables: Bloch vectors in phase space, Physical Review A94, 010301 (2016)
work page 2016
-
[40]
IQPopt: Fast optimization of instantaneous quantum polynomial circuits in JAX
E. Armengol and J. Bowles, Iqpopt: Fast optimization of instantaneous quantum polynomial circuits in jax, arXiv preprint arXiv:2501.04776 (2025)
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[41]
M. S. Rudolph, S. Lerch, S. Thanasilp, O. Kiss, O. Shaya, S. Vallecorsa, M. Grossi, and Z. Holmes, Trainability bar- riers and opportunities in quantum generative modeling, npj Quantum Information10, 116 (2024)
work page 2024
- [42]
- [43]
-
[44]
A note on the evaluation of generative models
L. Theis, A. v. d. Oord, and M. Bethge, A note on the evaluation of generative models, arXiv preprint arXiv:1511.01844 (2015)
work page internal anchor Pith review Pith/arXiv arXiv 2015
-
[45]
Q. Xu, G. Huang, Y. Yuan, C. Guo, Y. Sun, F. Wu, and K. Weinberger, An empirical study on evaluation met- rics of generative adversarial networks, arXiv preprint arXiv:1806.07755 (2018)
work page internal anchor Pith review Pith/arXiv arXiv 2018
-
[46]
A Practical Guide to Sample-based Statistical Distances for Evaluating Generative Models in Science
S. Bischoff, A. Darcher, M. Deistler, R. Gao, F. Gerken, M. Gloeckler, L. Haxel, J. Kapoor, J. K. Lappalainen, J. H. Macke,et al., A practical guide to sample-based statistical distances for evaluating generative models in science, arXiv preprint arXiv:2403.12636 (2024)
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[47]
S. C. Marshall, S. Aaronson, and V. Dunjko, Im- proved separation between quantum and classical com- puters for sampling and functional tasks, arXiv preprint arXiv:2410.20935 (2024)
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[48]
M. J. Bremner, R. Jozsa, and D. J. Shepherd, Classical simulation of commuting quantum computations implies collapse of the polynomial hierarchy, Proceedings of the Royal Society A: Mathematical, Physical and Engineer- ing Sciences467, 459 (2011)
work page 2011
-
[49]
M. J. Bremner, A. Montanaro, and D. J. Shepherd, Average-case complexity versus approximate simulation of commuting quantum computations, Physical review letters117, 080501 (2016)
work page 2016
-
[50]
N. Rahaman, A. Baratin, D. Arpit, F. Draxler, M. Lin, F. Hamprecht, Y. Bengio, and A. Courville, On the spec- tral bias of neural networks, inInternational Conference on Machine Learning(PMLR, 2019) pp. 5301–5310
work page 2019
-
[51]
Z.-Q. J. Xu, Y. Zhang, T. Luo, Y. Xiao, and Z. Ma, Frequency principle: Fourier analysis sheds light on deep neural networks, arXiv preprint arXiv:1901.06523 (2019)
work page internal anchor Pith review Pith/arXiv arXiv 1901
-
[52]
A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Sch¨ olkopf, and A. Smola, A kernel two-sample test, The Journal of Machine Learning Research13, 723 (2012)
work page 2012
- [53]
-
[54]
R. I. Kondor and J. Lafferty, Diffusion kernels on graphs and other discrete structures, inProceedings of the 19th international conference on machine learning, Vol. 2002 (2002) pp. 315–322
work page 2002
-
[55]
D. I. Shuman, S. K. Narang, P. Frossard, A. Ortega, and P. Vandergheynst, The emerging field of signal processing on graphs: Extending high-dimensional data analysis to networks and other irregular domains, IEEE signal pro- cessing magazine30, 83 (2013). 15
work page 2013
-
[56]
A. J. Smola and R. Kondor, Kernels and regulariza- tion on graphs, inLearning theory and kernel machines: 16th annual conference on learning theory and 7th kernel workshop, COLT/kernel 2003, Washington, DC, USA, august 24-27, 2003. Proceedings(Springer, 2003) pp. 144–158
work page 2003
-
[57]
Wu, The potts model, Reviews of modern physics 54, 235 (1982)
F.-Y. Wu, The potts model, Reviews of modern physics 54, 235 (1982)
work page 1982
-
[58]
S. Griffiths-Jones, A. Bateman, M. Marshall, A. Khanna, and S. R. Eddy, Rfam: an rna family database, Nucleic acids research31, 439 (2003)
work page 2003
-
[59]
D. L. Donoho and P. B. Stark, Uncertainty principles and signal recovery, SIAM Journal on Applied Mathematics 49, 906 (1989)
work page 1989
-
[60]
R. Meshulam, An uncertainty inequality for finite abelian groups, European Journal of Combinatorics27, 63 (2006)
work page 2006
- [61]
- [62]
-
[63]
Exponentially many initializations to avoid barren plateaus
A. Kulshrestha, R. Puig, D. Garc´ ıa-Mart´ ın, L. Cincio, I. Safro, Z. Holmes, and M. Cerezo, Exponentially many initializations to avoid barren plateaus, arXiv preprint arXiv:2606.18515 (2026)
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[64]
P.-y. Chiang, R. Ni, D. Y. Miller, A. Bansal, J. Geip- ing, M. Goldblum, and T. Goldstein, Loss landscapes are all you need: Neural network generalization can be ex- plained without the implicit bias of gradient descent, in The Eleventh International Conference on Learning Rep- resentations(2023). Appendix A: Proof of Proposition 1 Using the formU(θ) = (F...
work page 2023
-
[65]
The prefactor matrixJ Expanding the entryJ i,z, Ji,z = π d 2m i ·z−m i ·k i .(B9) The bilinear term 2m i ·zis the (i,z) entry of 2M Z T ∈ Z|O|×|Z|. The termm i ·k i depends only oniand is the i-th entry of (M⊙K)1 n, which we broadcast across|Z| columns by right-multiplying with1 T |Z|. This recovers (25): J= π d 2M ZT −((M⊙K)1 n)1 T |Z| .(B10)
-
[66]
The phase matrixE The matrixEis more involved. Using the explicit form (22) of the phase function Φ θ(z) = P g∈G θg ϕg(z), the entries decompose as Ei,z = X g∈G θg ϕg(z)−ϕ g(z⊖k i) .(B11) Splitting the gate setGinto subsetsG w of generators with weightw(i.e. exactlywnon-zero entries),G= F w Gw, the matrix decomposes additively asE= P w Ew, where [Ew]i,z =...
-
[67]
Total complexity For a gate set restricted to weight-wgenerators, build- ingE w requires assembling each of the 2 w pairs of ma- trices (B w σ ,C w σ) at costO(|G w|(|Z|+|O|)) each, plus the matrix eBw at costO(|G w||Z|). Performing the corresponding triple matrix multiplications then costs O(|Gw||O||Z|) per term. Adding the costO(|O||Z|n) of buildingJfro...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.