Pith. sign in

REVIEW 2 major objections 5 minor 67 references

Spectral Born machines learn integer data with a Fourier bias that may block overfitting even when heavily over-parameterized.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-10 23:36 UTC pith:OZVA5CLF

load-bearing objection Clean group-Fourier lift of IQP Born machines with solid classical estimators, real software, and a million-parameter RNA run; the overfitting-immunity claim is overstated relative to the evidence. the 2 major comments →

arxiv 2607.06675 v1 pith:OZVA5CLF submitted 2026-07-07 quant-ph

Spectral Born machines: classically trainable quantum generative models for discrete data

classification quant-ph
keywords spectral Born machinesquantum generative modelsgroup Fourier analysisHeisenberg-Weyl observablesgraph-spectral MMDtrain classical deploy quantuminteger dataoverparameterization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper introduces spectral Born machines: quantum generative models for vectors of integers that put a diagonal phase unitary between quantum Fourier transforms. The Fourier structure supplies an inductive bias toward smooth, low-order distributions on the integers, while the same structure lets training be done entirely classically by estimating Heisenberg-Weyl moments inside a graph-spectral maximum-mean-discrepancy loss. The authors release the procedure as a software module and show that gate sets matched to the bias reach the same accuracy with far fewer parameters than unstructured alternatives. They also train a 190-qubit, million-parameter model on scarce ribosomal-RNA sequences and find that the model matches low-order Fourier statistics only to the statistical floor of the data, suggesting that extreme over-parameterization need not produce memorization. If the classical hardness of sampling survives realistic parameter regimes, the approach yields a practical train-classical, deploy-quantum pipeline for discrete generative modeling.

Core claim

Spectral Born machines—Fourier-phase unitaries on qudits—can be trained at scale on classical hardware via batched Heisenberg-Weyl moment matching under a graph-spectral MMD, and their built-in low-order spectral bias appears to prevent overfitting even when the model has orders of magnitude more parameters than training samples.

What carries the argument

Fourier-phase unitaries (a diagonal phase layer D(θ) sandwiched between quantum Fourier transforms over Z_d^n) together with classical Monte-Carlo estimation of their Heisenberg-Weyl moments, which turn the graph-spectral MMD into an efficiently optimizable loss.

Load-bearing premise

That sampling the trained model stays classically hard once parameters are small and the distribution is spectrally concentrated—the regime the paper itself flags as open to possible dequantization.

What would settle it

Classically sample (or closely approximate) the output distribution of a trained spectral Born machine whose low-order Fourier coefficients match real data, using only the polynomial number of moments accessible during classical training; success would remove the quantum-deploy advantage.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Integer-structured generative tasks can be attacked with far smaller parameter counts by restricting the phase gates to low-degree Fourier generators.
  • Million-parameter quantum generative models become trainable today on classical GPUs, with quantum hardware needed only at inference.
  • Over-parameterization can be used aggressively for expressivity without the usual overfitting penalty when the model class is spectrally biased.
  • Graph-spectral MMD losses give a principled way to encode whether discrete variables are ordinal or purely categorical.
  • The same classical training loop can be scaled to distributed clusters, opening billion-parameter quantum generative models.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same Fourier-phase construction may supply a controllable spectral regularizer for other discrete generative families beyond Born machines.
  • If dequantization via low-order Fourier sparsity fails, spectral Born machines occupy a practically useful complexity middle ground where hardness is neither proved nor easily spoofed.
  • Matching the gate-set degrees to the cyclic distance of the target may become a standard inductive-bias design rule for any Fourier-based quantum generative model.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces spectral Born machines as a generalization of IQP Born machines to qudits via the group Fourier transform over Z_d^n. A Fourier-phase unitary (QFT layer, diagonal parameterized phase generated by Heisenberg-Weyl operators, inverse QFT) defines a generative model over integer vectors whose computational-basis distribution can be trained classically: Proposition 1 and the batched linear-algebra estimator (Apps. A–B) show that HW expectation values are Monte-Carlo estimable to inverse-polynomial additive error, enabling an unbiased graph-spectral MMD loss (App. C) based on cycle or complete graphs. The construction is released in the PennyLane tcdq module. Numerical experiments on a Potts model demonstrate that restricting gate degree yields comparable test MMD with far fewer parameters than the unrestricted gate set; a 190-qubit (95-qudit) model with >1 M parameters is trained on a scarce 93-nt rRNA data set and matches low-order Fourier coefficients to within the 1/√|X| floor. The authors interpret the latter as evidence that extreme over-parameterization may be immune to overfitting because of the built-in low-order spectral bias.

Significance. If the classical-trainability pipeline and the spectral inductive bias hold, the work supplies a practical route to large-scale quantum generative modeling for non-binary discrete data, removing the need for quantum gradient estimation and data loading during training. The clean Monte-Carlo theory (Prop. 1, Apps. A–C), the explicit batched estimator, and the public tcdq implementation are concrete engineering contributions that enable empirical exploration at the 100–1000 qubit scale. The Potts results give direct evidence that matching the model’s Fourier support to the data’s smoothness structure reduces parameter count without loss of performance. The rRNA experiment, while limited, is among the largest classically trained quantum generative models reported. These elements make the manuscript a useful step toward “train classical, deploy quantum” generative models; the ultimate quantum advantage remains conjectural pending hardness results in the trained, spectrally concentrated regime.

major comments (2)
  1. [abstract; Sec. V B; Fig. 5; Sec. VI A] The central interpretive claim that “highly over-parameterized spectral Born machines may be immune to overfitting, even in strongly data-scarce regimes” (abstract; Sec. VI A) is supported only by the rRNA experiment (Sec. V B). There a ~1.3 M-parameter model reaches MMD^{2} ≈ 0.0035 (comparable to the 1/√|X| ≈ 0.05 statistical floor) and its weight-1/2 Fourier coefficients visually match the empirical spectrum (Fig. 5). Because the training loss (Eq. 42) and the complete-graph heat kernel deliberately target only low-order Heisenberg-Weyl moments, agreement on those moments does not rule out (i) memorization of the 428 training sequences via higher-order correlations invisible to the loss, or (ii) collapse onto an incorrect but spectrally smooth distribution. No generative samples, high-order moment diagnostics, or likelihood proxies are reported. The claim should be substantially quali
  2. [abstract; Sec. IV; Sec. VI C] The abstract asserts that the models “remain classically hard to sample from in general.” Sec. IV (paragraph following the definition of spectral Born machines) and Sec. VI C correctly note that anticoncentration/random-circuit arguments do not apply once parameters are small and the distribution is spectrally concentrated, and that dequantization via low-order Fourier sparsity is an open risk. For the practically relevant trained models that are the point of the work, classical hardness is therefore conjectural. The unqualified phrasing in the abstract and introduction should be aligned with the more cautious discussion in Sec. VI C.
minor comments (5)
  1. [abstract; p. 1] Abstract and opening paragraph contain missing spaces (“We presentspectral”, “newtcdqmodule”).
  2. [Fig. 2; Fig. 4] Fig. 2 caption claims “similar training loss, suggesting o similar test performance”; the actual test curves appear only in Fig. 4. Cross-reference or merge the figures for clarity.
  3. [Sec. IV B; Sec. V] The heat-kernel bandwidth t (or equivalent mean operator weight) is a free hyper-parameter; a short sensitivity study or default-selection rule would help reproducibility.
  4. [Sec. V B] In the rRNA experiment the three-body gates are chosen by ranking empirical Fourier coefficients of the training set. This data-dependent gate selection should be stated more prominently as part of the model-construction pipeline.
  5. [throughout] Typographical inconsistencies: “Z n d” vs. “Z_d^n”, occasional missing punctuation after displayed equations, and “tcdq” sometimes rendered without backticks or italics.

Circularity Check

1 steps flagged

No load-bearing circularity; classical estimability and spectral MMD follow from QFT conjugation and graph Laplacians, with only minor self-citation of the prior TCDQ/IQP training paradigm.

specific steps
  1. self citation load bearing [Sec. I (TCDQ paragraph) and Sec. IV (definition of spectral Born machines)]
    "The possibility of training quantum generative models at scale in this way was first shown for instantaneous quantum polynomial (IQP) Born machines [12] … For the case d=2, spectral Born machines become the class of IQP Born machines and thus inherit similar complexity-theoretic classical hardness guarantees related to sampling [12, 43-45]."

    Reference [12] shares an author (Bowles) and supplies the classical-training paradigm that is reused. The citation is not load-bearing for the new Z_d^n mathematics (QFT conjugation, HW batch estimator, graph-spectral MMD), which is derived independently; it is therefore only a minor self-citation of the surrounding framework rather than a circular reduction of the paper's central claims.

full rationale

The paper's core technical claims are self-contained derivations. Proposition 1 and Appendix A obtain the Monte-Carlo estimator of Heisenberg-Weyl expectations directly from the Fourier-phase unitary form U=(F†)D( heta)F together with the conjugation FD(k,m)F†=D(m,-k) and unit-modulus averaging; nothing is fitted or defined in terms of the target loss. The graph-spectral MMD (Eqs. 33-42) is constructed from the Laplacian of the chosen product graph (cycle or complete) and reduces to a weighted mixture of |<D(k,0)>|^{2} differences by the standard spectral theorem for circulant matrices; the bandwidth t and gate-set restrictions are free design choices, not circular redefinitions of the data. Spectral bias (Sec. IV A, Eqs. 28-29) is an immediate consequence of low-weight generators producing low-frequency phase functions. The rRNA and Potts experiments report empirical MMD values and Fourier visualizations; they do not fit a parameter on one quantity and then re-label a related quantity as a prediction. Self-citations to the authors' prior IQP/TCDQ work supply the overall train-classical-deploy-quantum template but are not used to force uniqueness or to smuggle an ansatz that itself lacks independent support. Concurrent work [30] is acknowledged as complementary. Consequently the derivation chain does not reduce any claimed result to its own inputs by construction.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 2 invented entities

The central claims rest on standard group Fourier analysis, the known classical hardness of IQP sampling (inherited), the Monte-Carlo estimability of phase differences, and several free design choices (gate-set restrictions, heat-kernel bandwidth, initialization scales). No new physical entities are postulated; the model class itself is the invented construct.

free parameters (5)
  • heat-kernel bandwidth t (or equivalent mean operator weight)
    Chosen so that the average weight of sampled Fourier modes is 3 (Potts) or 2 (RNA); controls the low-pass character of the MMD and therefore which statistics the model is forced to match.
  • gate-set degree / weight cut-offs and selected three-body gates
    Degree-1/2/3 versus all-degree for Potts; all two-body plus magnitude-selected three-body for RNA. Directly determines parameter count and spectral bias.
  • initialization standard deviations for multi-qudit gates
    0.001 (Potts), 0.01 and 0.0001 (RNA models); chosen by hand to keep initial correlations short-range.
  • Monte-Carlo batch sizes |K| and |Z|
    1000 (Potts) or 500 (RNA); control gradient noise and estimator variance.
  • learning rate and iteration count
    Initial LR 0.001, 10 000 iterations (Potts); analogous schedule for RNA. Standard optimizer free parameters.
axioms (4)
  • domain assumption Sampling from IQP circuits (and by extension spectral Born machines at d=2) is classically hard under standard complexity assumptions
    Invoked in Sec. IV to motivate quantum advantage; inherited from Bremner et al. and subsequent IQP literature.
  • domain assumption Phase function Phi_theta(z) is classically computable in poly(n,d) time for the chosen gate sets
    Required for Proposition 1; satisfied by the explicit product-of-cosines form of the Heisenberg-Weyl generators.
  • domain assumption Graph Laplacian heat kernel on cycle or complete graphs supplies a suitable notion of closeness for ordinal/categorical discrete data
    Sec. IV B; standard in graph signal processing but an modeling choice for the generative loss.
  • ad hoc to paper Low-order Fourier concentration is a desirable inductive bias that forbids memorization of sparse empirical distributions
    Sec. VI A; motivated by uncertainty principles and neural-network spectral bias literature, but not proved for the present model class.
invented entities (2)
  • spectral Born machine (Fourier-phase unitary generative model over Z_d^n) no independent evidence
    purpose: Provide a classically trainable, spectrally biased quantum generative model for integer-structured data
    Defined in Sec. IV as the computational-basis distribution of a Fourier-phase circuit; the central object of the paper.
  • tcdq PennyLane module independent evidence
    purpose: Ship the classical training estimators so others can reproduce and extend the experiments
    Released as part of PennyLane Labs; software artifact rather than a physical entity.

pith-pipeline@v1.1.0-grok45 · 27907 in / 3505 out tokens · 48775 ms · 2026-07-10T23:36:17.460683+00:00 · methodology

0 comments
read the original abstract

We present \emph{spectral Born machines}, a class of quantum generative models that results from viewing and generalizing the class of IQP Born machines through the lens of group Fourier analysis. These quantum models exploit the quantum Fourier transform to create an inductive bias that make them naturally suited to learning integer-structured data, while remaining classically hard to sample from in general. Similar to IQP Born machines, spectral Born machines can be trained efficiently at scale on classical hardware via a maximum mean discrepancy loss based on graph spectral analysis, which we make available in a new \emph{tcdq} module of the PennyLane software platform. In numerical experiments, we show how the spectral bias of the model leads to significantly reduced parameter counts compared to unstructured approaches, and demonstrate the scalability of the software by training a 190-qubit model with over 1 million parameters to successfully learn a distribution of 93 nucleotide-long ribosomal RNA. Our results suggest that highly over-parameterized spectral Born machines may be immune to overfitting, even in strongly data-scarce regimes.

Figures

Figures reproduced from arXiv: 2607.06675 by Austin Huang, Evan Peters, Jason Pye, Joseph Bowles, Soran Jahangiri, Vasilis Belis, William Maxwell.

Figure 1
Figure 1. Figure 1: FIG. 1. (a) A spectral Born machine consists of a diagonal parameterized unitary sandwiched between two layers of quantum [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2. The four spectral Born machines (with varying parameter counts) used for Potts model experiments achieve similar [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3. Example configurations from the Potts model on a [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4. The inductive bias of the spectral Born machine allows small models to achieve comparable performance compared to [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: FIG. 5. Comparing visualizations of the Fourier coefficients of the complete dataset (left) versus trained spectral Born machine [PITH_FULL_IMAGE:figures/full_fig_p010_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: FIG. 6. A spectral Born machine model with [PITH_FULL_IMAGE:figures/full_fig_p011_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

67 extracted references · 67 canonical work pages · 18 internal anchors

  1. [1]

    of the complex-valued functionf:G→Cand its inverse are given by ˆf(k) := X x∈G f(x)χ k(x), f(x) = 1 |G| X k∈G ˆf(k)χ ∗ k(x),(1) where the group charactersχ k :G→Ccorrespond to the one-dimensional irreducible representations ofG. We work with the product cyclic groupG=Z n d, which nat- urally relates to vectors of integers, for which the char- acters are g...

  2. [2]

    can be realised via a tensor product of quantum Fourier transforms ofZ d. Concretely, the QFT overZ d and its inverse are F|x⟩= 1√ d X k ωkx |k⟩,(2) F † |x⟩= 1√ d X k ω−kx |k⟩,(3) so that the quantum Fourier transform overZ n d and its inverse are F ⊗n |x⟩= 1√ dn X k∈Zn d ωk·x |k⟩,(4) (F †)⊗n |x⟩= 1√ dn X k∈Zn d ω−k·x |k⟩.(5) Note that ford= 2, we haveF=H...

  3. [3]

    Shift-invariant graphs: cycleC d and completeK d When eachG i is invariant under cyclic shifts of labels, which is the case for the cycle graph or the complete graph, the LaplacianL i is a circulant matrix and the eigenvectorsu k of the Laplacian are the characters of Zn d, e⊤ x uk = 1√ dn ωk·x.(39) Since the kernelK=f(L) is also diagonalised by this basi...

  4. [4]

    how much data we would need to train models of increasing size while keeping overfitting under control

    Finite-sample estimation Here we show how to efficiently estimate and train the MMD2 loss. Definep X to be an empirical distribution of the datasetXsampled from the ground truth distribu- tionp pX (x) = 1 |X | X y∈X I[y=x],(46) withIthe indicator function. Note that⟨D(k,0)⟩ pX can be computed exactly and efficiently sincep X has sparse support. Let \⟨D(k,...

  5. [5]

    Scaling Laws for Neural Language Models

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, Scaling laws for neural language models, arXiv preprint arXiv:2001.08361 (2020)

  6. [6]

    A. G. Wilson, Deep learning is not so mysterious or dif- ferent, arXiv preprint arXiv:2503.02113 (2025)

  7. [7]

    A. G. Wilson and P. Izmailov, Bayesian deep learning and a probabilistic perspective of generalization, Advances in neural information processing systems33, 4697 (2020)

  8. [8]

    C. Liu, L. Zhu, and M. Belkin, Loss landscapes and op- timization in over-parameterized non-linear systems and neural networks, Applied and Computational Harmonic Analysis59, 85 (2022)

  9. [9]

    Hoefler, T

    T. Hoefler, T. H¨ aner, and M. Troyer, Disentangling hype from practicality: On realistically achieving quantum ad- vantage, Communications of the ACM66, 82 (2023)

  10. [10]

    Bowles, D

    J. Bowles, D. Wierichs, and C.-Y. Park, Backpropagation scaling in parameterised quantum circuits, Quantum9, 1873 (2025)

  11. [11]

    Abbas, R

    A. Abbas, R. King, H.-Y. Huang, W. J. Huggins, R. Movassagh, D. Gilboa, and J. McClean, On quan- tum backpropagation, information reuse, and cheating measurement collapse, Advances in Neural Information Processing Systems36, 44792 (2023)

  12. [12]

    Coyle, S

    B. Coyle, S. Raj, N. Mathur, E. A. Cherrat, N. Jain, S. Kazdaghli, and I. Kerenidis, Training-efficient density quantum machine learning, npj Quantum Information 11, 172 (2025)

  13. [13]

    Chinzei, S

    K. Chinzei, S. Yamano, Q. H. Tran, Y. Endo, and H. Os- hima, Trade-off between gradient measurement efficiency and expressivity in deep quantum neural networks, npj Quantum Information11, 79 (2025)

  14. [14]

    Scalable On-Hardware Training of Quantum Neural Networks and Application to Clinical Data Imputation

    N. Mathur, P. K. Barkoutsos, M. Yamada, M. Roetteler, and I. Kerenidis, Scalable on-hardware training of quan- tum neural networks and application to clinical data im- putation, arXiv preprint arXiv:2606.03517 (2026)

  15. [15]

    Adaptive directional gradients for parameterised quantum circuits

    B. Coyle, S. Raj, V. Umathe, E. A. Cher- rat, and E. Kashefi, Adaptive directional gradients for parameterised quantum circuits, arXiv preprint arXiv:2606.09734 (2026)

  16. [16]

    Recio-Armengol, S

    E. Recio-Armengol, S. Ahmed, and J. Bowles, Train on classical, deploy on quantum: scaling generative quantum machine learning to a thousand qubits, arXiv preprint arXiv:2503.02934 (2025)

  17. [17]

    Kasture, O

    S. Kasture, O. Kyriienko, and V. E. Elfving, Protocols for classically training quantum generative models on probability distributions, Physical Review A108, 042406 (2023)

  18. [18]

    Bak´ o, Z

    B. Bak´ o, Z. Kolarovszki, and Z. Zimbor´ as, Fermionic born machines: Classical training of quantum genera- tive models based on fermion sampling, arXiv preprint arXiv:2511.13844 (2025)

  19. [19]

    Efficient training of photonic quantum generative models

    F. Gottlieb, R. Mezher, B. Ventura, S. Mansfield, and A. Salavrakos, Efficient training of photonic quan- tum generative models, arXiv preprint arXiv:2603.08793 (2026)

  20. [20]

    Kolarovszki, B

    Z. Kolarovszki, B. Bak´ o, M. Oszmaniec, C. Oh, and Z. Zimbor´ as, Generative modeling with gaussian boson sampling: classically trainable bosonic born machines, arXiv preprint arXiv:2603.11195 (2026)

  21. [21]

    Kurkin, U

    A. Kurkin, U. Chabaud, Z. Kolarovszki, B. Bak´ o, Z. Zim- bor´ as, and V. Dunjko, Universality of classically train- able, quantum-deployed boson-sampling generative mod- els, arXiv preprint arXiv:2603.11014 (2026)

  22. [22]

    Quantum Fourier Generative Models Trainable at Large Scale

    C. T¨ uys¨ uz, O. Kyriienko, and M. Grossi, Quantum fourier generative models trainable at large scale, arXiv preprint arXiv:2606.28483 (2026)

  23. [23]

    Herrero-Gonzalez, B

    M. Herrero-Gonzalez, B. Coyle, K. McDowall, R. Grassie, S. Beentjes, A. Khamseh, and E. Kashefi, The born ul- timatum: Conditions for classical surrogation of quan- 14 tum generative models with correlators, arXiv preprint arXiv:2511.01845 (2025)

  24. [24]

    Ball´ o-Gimbernat, M

    O. Ball´ o-Gimbernat, M. Arroyo-S´ anchez, P. Garc´ ıa- Molina, A. Garriga, and F. Vilari˜ no, Shallow instanta- neous quantum polynomial-time circuits for generative modeling on noisy intermediate-scale quantum hardware, Physical Review A113, 042617 (2026)

  25. [25]

    Herbst, I

    S. Herbst, I. Brandi´ c, and A. P´ erez-Salinas, Limits of quantum generative models with classical sampling hard- ness, arXiv preprint arXiv:2512.24801 (2025)

  26. [26]

    J. Slim, S. Monaco, F. Rehm, D. Kr¨ ucker, and K. Borras, An iqp born machine for calorimeter image generation at 64 qubits with compiled-iqp deployment, arXiv preprint arXiv:2605.27735 (2026)

  27. [27]

    Parity Supervision as a Driver of Generalization in Quantum Generative Modeling

    M. Baumann, D. Hein, S. Udluft, T. Rohe, C. Linnhoff- Popien, and J. Stein, Parity supervision as a driver of generalization in quantum generative modeling, arXiv preprint arXiv:2605.10258 (2026)

  28. [28]

    Rosca, T

    M. Rosca, T. Weber, A. Gretton, and S. Mohamed, A case for new neural network smoothness constraints, PMLR137, 21 (2020)

  29. [29]

    Spectral methods: crucial for machine learning, natural for quantum computers?

    V. Belis, J. Bowles, R. Gupta, E. Peters, and M. Schuld, Spectral methods: crucial for machine learn- ing, natural for quantum computers?, arXiv preprint arXiv:2603.24654 (2026)

  30. [30]

    Dherin, M

    B. Dherin, M. Munn, M. Rosca, and D. Barrett, Why neural networks find simple solutions: The many regu- larizers of geometric complexity, Advances in Neural In- formation Processing Systems35, 2333 (2022)

  31. [31]

    Seehttps://docs

    The tcdq module is currently under development but available as part of PennyLane labs. Seehttps://docs. pennylane.ai/en/latest/code/qp_labs.html

  32. [32]

    PennyLane: Automatic differentiation of hybrid quantum-classical computations

    V. Bergholm, J. Izaac, M. Schuld, C. Gogolin, S. Ahmed, V. Ajith, M. S. Alam, G. Alonso-Linaje, B. Akash- Narayanan, A. Asadi,et al., Pennylane: Automatic dif- ferentiation of hybrid quantum-classical computations, arXiv preprint arXiv:1811.04968 (2018)

  33. [33]

    Bradbury, R

    J. Bradbury, R. Frostig, P. Hawkins, M. J. John- son, Y. Katariya, C. Leary, D. Maclaurin, G. Nec- ula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang, JAX: composable transformations of Python+NumPy programs (2018)

  34. [34]

    R. J. Banks, A. Crippa, M. Traube, J. Unger, C. Ertler, and W. Lechner, Qudit extension of parameterized iqp circuits: A generative quantum machine learning ap- proach to integer data, arXiv preprint arXiv:2606.28236 (2026)

  35. [35]

    Dherin, M

    B. Dherin, M. Munn, M. Rosca, and D. G. Barrett, Why neural networks find simple solutions: The many reg- ularizers of geometric complexity, inAdvances in Neu- ral Information Processing Systems, edited by A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho (2022)

  36. [36]

    Krebsbach, F

    M. Krebsbach, F. Reiter, T. Wellens, H.-H. Kowal- ski, and A. Abedi, Encoding numerical data for generative quantum machine learning, arXiv preprint arXiv:2603.23407 (2026)

  37. [37]

    I. R. Kondor,Group theoretical methods in machine learning(Columbia University, 2008)

  38. [38]

    A. M. Childs and W. Van Dam, Quantum algorithms for algebraic problems, Reviews of Modern Physics82, 1 (2010)

  39. [39]

    Asadian, P

    A. Asadian, P. Erker, M. Huber, and C. Kl¨ ockl, Heisenberg-weyl observables: Bloch vectors in phase space, Physical Review A94, 010301 (2016)

  40. [40]

    IQPopt: Fast optimization of instantaneous quantum polynomial circuits in JAX

    E. Armengol and J. Bowles, Iqpopt: Fast optimization of instantaneous quantum polynomial circuits in jax, arXiv preprint arXiv:2501.04776 (2025)

  41. [41]

    M. S. Rudolph, S. Lerch, S. Thanasilp, O. Kiss, O. Shaya, S. Vallecorsa, M. Grossi, and Z. Holmes, Trainability bar- riers and opportunities in quantum generative modeling, npj Quantum Information10, 116 (2024)

  42. [42]

    Liu and L

    J.-G. Liu and L. Wang, Differentiable learning of quan- tum circuit born machines, Physical Review A98, 062324 (2018)

  43. [43]

    Kurkin, K

    A. Kurkin, K. Shen, S. Pielawa, H. Wang, and V. Dun- jko, Universality and kernel-adaptive training for clas- sically trained, quantum-deployed generative models, arXiv preprint arXiv:2510.08476 (2025)

  44. [44]

    A note on the evaluation of generative models

    L. Theis, A. v. d. Oord, and M. Bethge, A note on the evaluation of generative models, arXiv preprint arXiv:1511.01844 (2015)

  45. [45]

    Q. Xu, G. Huang, Y. Yuan, C. Guo, Y. Sun, F. Wu, and K. Weinberger, An empirical study on evaluation met- rics of generative adversarial networks, arXiv preprint arXiv:1806.07755 (2018)

  46. [46]

    A Practical Guide to Sample-based Statistical Distances for Evaluating Generative Models in Science

    S. Bischoff, A. Darcher, M. Deistler, R. Gao, F. Gerken, M. Gloeckler, L. Haxel, J. Kapoor, J. K. Lappalainen, J. H. Macke,et al., A practical guide to sample-based statistical distances for evaluating generative models in science, arXiv preprint arXiv:2403.12636 (2024)

  47. [47]

    S. C. Marshall, S. Aaronson, and V. Dunjko, Im- proved separation between quantum and classical com- puters for sampling and functional tasks, arXiv preprint arXiv:2410.20935 (2024)

  48. [48]

    M. J. Bremner, R. Jozsa, and D. J. Shepherd, Classical simulation of commuting quantum computations implies collapse of the polynomial hierarchy, Proceedings of the Royal Society A: Mathematical, Physical and Engineer- ing Sciences467, 459 (2011)

  49. [49]

    M. J. Bremner, A. Montanaro, and D. J. Shepherd, Average-case complexity versus approximate simulation of commuting quantum computations, Physical review letters117, 080501 (2016)

  50. [50]

    Rahaman, A

    N. Rahaman, A. Baratin, D. Arpit, F. Draxler, M. Lin, F. Hamprecht, Y. Bengio, and A. Courville, On the spec- tral bias of neural networks, inInternational Conference on Machine Learning(PMLR, 2019) pp. 5301–5310

  51. [51]

    Z.-Q. J. Xu, Y. Zhang, T. Luo, Y. Xiao, and Z. Ma, Frequency principle: Fourier analysis sheds light on deep neural networks, arXiv preprint arXiv:1901.06523 (2019)

  52. [52]

    Gretton, K

    A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Sch¨ olkopf, and A. Smola, A kernel two-sample test, The Journal of Machine Learning Research13, 723 (2012)

  53. [53]

    Li, W.-C

    C.-L. Li, W.-C. Chang, Y. Cheng, Y. Yang, and B. P´ oczos, Mmd gan: Towards deeper understanding of moment matching network, Advances in Neural Informa- tion Processing Systems30(2017)

  54. [54]

    R. I. Kondor and J. Lafferty, Diffusion kernels on graphs and other discrete structures, inProceedings of the 19th international conference on machine learning, Vol. 2002 (2002) pp. 315–322

  55. [55]

    D. I. Shuman, S. K. Narang, P. Frossard, A. Ortega, and P. Vandergheynst, The emerging field of signal processing on graphs: Extending high-dimensional data analysis to networks and other irregular domains, IEEE signal pro- cessing magazine30, 83 (2013). 15

  56. [56]

    A. J. Smola and R. Kondor, Kernels and regulariza- tion on graphs, inLearning theory and kernel machines: 16th annual conference on learning theory and 7th kernel workshop, COLT/kernel 2003, Washington, DC, USA, august 24-27, 2003. Proceedings(Springer, 2003) pp. 144–158

  57. [57]

    Wu, The potts model, Reviews of modern physics 54, 235 (1982)

    F.-Y. Wu, The potts model, Reviews of modern physics 54, 235 (1982)

  58. [58]

    Griffiths-Jones, A

    S. Griffiths-Jones, A. Bateman, M. Marshall, A. Khanna, and S. R. Eddy, Rfam: an rna family database, Nucleic acids research31, 439 (2003)

  59. [59]

    D. L. Donoho and P. B. Stark, Uncertainty principles and signal recovery, SIAM Journal on Applied Mathematics 49, 906 (1989)

  60. [60]

    Meshulam, An uncertainty inequality for finite abelian groups, European Journal of Combinatorics27, 63 (2006)

    R. Meshulam, An uncertainty inequality for finite abelian groups, European Journal of Combinatorics27, 63 (2006)

  61. [61]

    K. Shen, S. Pielawa, V. Dunjko, and H. Wang, Character- izing trainability of instantaneous quantum polynomial circuit born machines, arXiv preprint arXiv:2602.11042 (2026)

  62. [62]

    Lerch, J

    S. Lerch, J. Bowles, R. Puig, E. Armengol, Z. Holmes, and S. Thanasilp, Iqp born machines under data- dependent and agnostic initialization strategies, arXiv preprint arXiv:2603.14576 (2026)

  63. [63]

    Exponentially many initializations to avoid barren plateaus

    A. Kulshrestha, R. Puig, D. Garc´ ıa-Mart´ ın, L. Cincio, I. Safro, Z. Holmes, and M. Cerezo, Exponentially many initializations to avoid barren plateaus, arXiv preprint arXiv:2606.18515 (2026)

  64. [64]

    Chiang, R

    P.-y. Chiang, R. Ni, D. Y. Miller, A. Bansal, J. Geip- ing, M. Goldblum, and T. Goldstein, Loss landscapes are all you need: Neural network generalization can be ex- plained without the implicit bias of gradient descent, in The Eleventh International Conference on Learning Rep- resentations(2023). Appendix A: Proof of Proposition 1 Using the formU(θ) = (F...

  65. [65]

    The termm i ·k i depends only oniand is the i-th entry of (M⊙K)1 n, which we broadcast across|Z| columns by right-multiplying with1 T |Z|

    The prefactor matrixJ Expanding the entryJ i,z, Ji,z = π d 2m i ·z−m i ·k i .(B9) The bilinear term 2m i ·zis the (i,z) entry of 2M Z T ∈ Z|O|×|Z|. The termm i ·k i depends only oniand is the i-th entry of (M⊙K)1 n, which we broadcast across|Z| columns by right-multiplying with1 T |Z|. This recovers (25): J= π d 2M ZT −((M⊙K)1 n)1 T |Z| .(B10)

  66. [66]

    The phase matrixE The matrixEis more involved. Using the explicit form (22) of the phase function Φ θ(z) = P g∈G θg ϕg(z), the entries decompose as Ei,z = X g∈G θg ϕg(z)−ϕ g(z⊖k i) .(B11) Splitting the gate setGinto subsetsG w of generators with weightw(i.e. exactlywnon-zero entries),G= F w Gw, the matrix decomposes additively asE= P w Ew, where [Ew]i,z =...

  67. [67]

    remove the diagonal

    Total complexity For a gate set restricted to weight-wgenerators, build- ingE w requires assembling each of the 2 w pairs of ma- trices (B w σ ,C w σ) at costO(|G w|(|Z|+|O|)) each, plus the matrix eBw at costO(|G w||Z|). Performing the corresponding triple matrix multiplications then costs O(|Gw||O||Z|) per term. Adding the costO(|O||Z|n) of buildingJfro...