Pith. sign in

REVIEW 3 major objections 5 minor 67 references

Training with maximally entangled quantum data exponentially shrinks the loss improvement an optimizer can find in a local neighborhood.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 18:05 UTC pith:AMVW4GUQ

load-bearing objection Real new theorem with a clean proof in PU(d); the experimental bridge to PQC training is qualitative, not a tight confirmation, and one proof line in Theorem 1(iii) needs a small correction. the 3 major comments →

arxiv 2509.10141 v1 pith:AMVW4GUQ submitted 2025-09-12 quant-ph

Loss Behavior in Supervised Learning with Entangled States

classification quant-ph MSC 81P6881P4568Q12 PACS 03.67.-a03.67.Lx
keywords quantum machine learningsupervised learningentangled training dataloss landscapebarren plateausparameterized quantum circuitstrainabilityentanglement entropy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Quantum machine learning can learn a unitary operator from a single maximally entangled training sample, and earlier work showed such samples reduce the risk after training. This paper asks what that benefit costs during training. It proves that for models that can explore all unitaries, the best possible loss improvement inside a fixed-radius neighborhood is exponentially smaller when the training sample is maximally entangled than when it is separable, and the distance to a global minimum is exponentially larger. Simulations with parameterized circuits confirm the effect grows with circuit expressivity and show that entanglement entropy, not Schmidt rank, is the quantity that predicts trainability.

Core claim

The central claim is Theorem 2: for a target U in U(d) with d = 2^n, if two intermediate solutions Vψ and VΦ have equal loss L with respect to a separable and a maximally entangled training sample, then for any radius R up to the value that guarantees a zero-loss solution for the separable sample, the ratio of achievable loss improvements satisfies ΛΦ(U,VΦ,R)/Λψ(U,Vψ,R) ∈ O(1/2^{n/2}). So in a neighborhood where the separable sample can reach a perfect solution, the best improvement available from the maximally entangled sample is exponentially small. Theorem 1 adds that the distance to a global minimum is exponentially larger for the entangled sample unless the current solution is already e

What carries the argument

The argument runs on the metric d'_F on the projective unitary group PU(d), the Frobenius-norm distance after optimizing over global phases, which for a maximally entangled sample equals sqrt(2d) times sqrt(1 − sqrt(F_{U,Φ}(V))) and therefore ties the loss directly to the distance from the target operator. Lemma 2 bounds the maximal fidelity achievable inside a ball B(V,R) in this metric for separable and maximally entangled samples; the improvement Λα(U,V,R) then reduces to sin(2γ − β) sin(β). The exponential separation comes from sin(β_ent) ≤ R/√d, so the improvement for the maximally entangled sample has a dimension factor in the denominator that the separable sample does not.

Load-bearing premise

The load-bearing premise is that optimizing over the full unitary group with Frobenius-norm balls faithfully captures the loss-landscape geometry of practical PQCs; the paper itself notes in its conclusion that the appropriate metric depends on the PQC, and the analytical bounds are only numerically validated for n = 5 qubits and four ansatz families.

What would settle it

Simulate supervised learning of a random n-qubit target with a highly expressive ansatz that can exactly represent the target, for increasing n, and measure the ratio of the best loss improvement in a ball of radius R for a maximally entangled sample versus a separable sample, taking R to be the distance where the separable sample first reaches zero loss. If that ratio does not shrink as 1/2^{n/2}, the exponential bound in Theorem 2 fails for that circuit family.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • For highly expressive models, using maximally entangled training samples makes gradients and loss differences exponentially small in local neighborhoods, which should make optimization substantially harder as the number of qubits grows.
  • The distance to a global minimum can be exponentially larger for the entangled sample, so an optimizer that starts far from the target may fail to find any improvement at all.
  • The detrimental effect grows with PQC expressivity: circuits that can explore more of the unitary group are more susceptible to entanglement-induced loss concentration, while circuits that already contain the target structure can avoid it.
  • Non-maximally entangled states offer a practical middle ground: high Schmidt rank keeps the risk low, while low entanglement entropy preserves trainability, suggesting warm-starting with low-entropy samples and fine-tuning with high-entropy ones.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If gradient magnitude tracks the local loss improvement, Theorem 2 implies a concrete no-free-lunch-style trade-off: the known sample-efficiency benefit of entangled data is paid for in optimization hardness, which may limit the practical quantum advantage of such schemes.
  • Because the proof uses the Frobenius metric on the full unitary group, the exponential bound should be read as a statement about a family of loss landscapes; for PQCs whose parameter-space metric differs, the effect could be weaker or stronger than the bound.
  • The experiments' finding that entanglement entropy predicts trainability while Schmidt rank does not suggests a practical pre-screening rule for training states: choose NME states with high Schmidt rank but concentrated Schmidt coefficients.
  • A direct testable extension is to measure gradient norms for n-qubit PQCs under separable versus maximally entangled training and check whether the ratio decays as 2^{-n/2}; this would connect the loss-improvement bound to the more standard barren-plateau diagnostics.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies how entanglement in the training data affects the trainability of supervised quantum models that learn a unitary operator. The analytical core is set in the idealized space PU(d) with the global-phase-corrected Frobenius metric d'_F. For a target U, a current hypothesis V, and a maximally entangled training state |Φ>, the loss L_{U,Φ}(V) is shown to be directly tied to d'_F(U,V) via Eq. (23)-(25). Lemma 1 gives the exact minimal d'_F-distance from V to any operator W achieving a prescribed fidelity f_W for a separable training state; Lemma 3 gives the analogous lower bound for a maximally entangled state. Theorem 1 compares distances to zero-loss optima: the distance is exponentially larger for the entangled sample except when V is already exponentially close to U. Theorem 2 shows that, for intermediate solutions with equal loss L and for any Frobenius ball of radius R at most the separable zero-loss radius, the best achievable loss improvement with a maximally entangled sample is O(1/2^{n/2}) times that for a separable sample. The paper then reports numerical constrained-optimization experiments on 5-qubit PQCs with four ansatzes and varying layer counts, together with experiments using non-maximally entangled states. The authors conclude that highly expressive PQCs are more susceptible to loss concentration induced by entangled training data, and that entanglement entropy, rather than Schmidt rank, is the better predictor of trainability.

Significance. If the PU(d)-level claim is taken as the main result, this is a valuable, self-contained contribution. The derivation is constructive: Lemma 1 explicitly builds the unitary that attains the distance bound, and Theorem 2 provides a rigorous asymptotic ratio without fitted parameters. The paper also ships reproducible simulation code and data (repository [49]), and it makes falsifiable predictions about the ordering of improvements for separable, maximally entangled, and non-maximally entangled samples. The entanglement-entropy ordering in Figures 7-8 is a useful empirical observation that goes beyond existing barren-plateau results. The main caveat is that the analytical result lives in the metric space (PU(d), d'_F), while the experimental support is obtained from 2-norm balls in PQC parameter space; the paper explicitly acknowledges this gap but does not close it. The strength of the practical claim -- that maximally entangled training samples exponentially limit trainability of PQCs -- is therefore weaker than the analytical result. This is a resolvable scoping issue, but it is load-bearing for the paper's broader narrative.

major comments (3)
  1. [§5.1-5.2, Eq. (45) and Theorem 2] The experimental protocol replaces the Frobenius balls B(V,R) of Theorem 2 by parameter-space balls B(θ0,R) = {θ : ||θ-θ0||_2 ≤ R}, but no relation is established between the radius R in the two spaces. For the deep ansatzes used, p is as large as 176, so the image of a 2-norm ball of radius R can contain unitaries whose d'_F-diameter is much larger than R; first-order estimates give ~√p·R. Thus R_max = 4 in Figure 4 does not confine the optimizer to the Frobenius ball of radius R_sep ≤ 2 in which Theorem 2 applies. The small Λ_Φ observed in Figure 6 could therefore be an artifact of constrained SLSQP failing inside the parameter ball, rather than a geometric absence of low-loss unitaries in the corresponding Frobenius ball. Moreover, Theorem 2 assumes Vψ and VΦ have equal loss L, whereas the experiments use the same starting point θ0 for both samples, so the initial losses generally dif
  2. [§4.1 / Appendix A, Theorem 1(iii)] The proof of Theorem 1(iii) derives γ ≤ O(1/√d) from Eq. (152) and then writes sin²γ ≤ γ, yielding O(1/2^{n/2}). The theorem statement claims O(1/2^n). The stronger bound follows immediately from sin²γ ≤ γ², but as written the proof does not establish the stated result. In addition, Eq. (154) contains a typographical error: it should read L_{U,Φ}(V), not L_{U†V,Φ}, and the angle in the preceding line should be γ_{U,Φ}(V), not γ_{U,ψ}(V). This is a local fix, but it should be corrected because Theorem 1(iii) is one of the paper's formal claims.
  3. [§4.2 / Eq. (42)-(44)] Theorem 2 compares improvements at equal loss L and equal Frobenius radius R. The upper bound in Eq. (193) and the subsequent bound on sin(β_ent) ≤ R/√d do show that, for a fixed radius, the entangled improvement is exponentially suppressed. However, the theorem does not by itself imply that a typical optimization trajectory starting from the same θ0 will experience this suppression, because the starting losses L_ψ(V(θ0)) and L_Φ(V(θ0)) are generally different. The experimental evaluation in Section 5.2 does not condition on equal starting loss; it instead chooses R_max as the distance at which the separable sample reaches zero loss. This is a reasonable heuristic, but it does not directly instantiate Theorem 2. The claim should be qualified accordingly, or an additional result should be supplied for unequal starting losses.
minor comments (5)
  1. [§2.4] Typo: 'the the Kullback-Leibler divergence'.
  2. [Appendix A, Eq. (154)] As noted in the major comments, L_{U†V,Φ} should be L_{U,Φ}(V), and the variable γ_{U,ψ}(V) in the preceding display should be γ_{U,Φ}(V).
  3. [Figure 5] The caption says 'distance to the closest minimum' but the text defines it as the smallest R such that L_{U,ψ}(V(θ)) ≤ 10^{-3}; the wording should be aligned with the operative definition.
  4. [§5.2.1, no-entanglement discussion] The explanation that no-entanglement is universal for single-qubit rotations is clear, but it would help to state explicitly that this is the reason the entries for l=16 in Figure 6 are an exception rather than evidence against the trend.
  5. [Appendix C, Eq. (224)] The notation P_exp and P_Haar is introduced, but the bins are defined only implicitly; please make the bin partition explicit in the equation or in the preceding sentence.

Circularity Check

0 steps flagged

No circularity: the exponential bounds follow from explicit geometric lemmas and external inequalities; self-citations are background and Section 8 honestly flags the PQC-metric caveat.

full rationale

The central derivation chain is self-contained. Lemma 1 (Section A) proves the constant-distance zero-fidelity result for separable states by deriving an upper bound from the external diagonal-element bound of Tromborg-Waldenström [56] and then explicitly constructing a unitary that attains it; it is not assumed. Lemma 3 derives the dimension-dependent bound for maximally entangled states from the operator decomposition and Cauchy-Schwarz. Theorems 1 and 2 are algebraic consequences of these lemmas plus standard inequalities [59,60]; no constant is fitted to data and no 'prediction' is a renamed input. The identity d'_F(U,V)=sqrt(2d(1-sqrt(F_{U,Phi}(V)))) (Eqs. 23-25) is used as a tool, but the comparison between separable and entangled samples is a genuine two-sided geometric statement, not a tautology. Self-citations such as [17] are cited for background risk results (Eq. 7 from [15]) and are not load-bearing for the trainability theorems. The paper explicitly states the idealization of optimizing over all of PU(d) and acknowledges in Section 8 that the Frobenius metric may not capture PQC parameter-space geometry; the experimental parameter-space balls (Eq. 45) differ from the theoretical Frobenius balls. That is an external-validity limitation, flagged by the authors themselves, not a circular reduction of the theorem to its assumptions. Accordingly no circular step is present.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The analytical derivation uses standard unitary-group geometry and a cited linear-algebra bound; no free parameters are fitted. The idealization of optimizing over all unitaries is an explicit domain assumption, not a hidden input. No new entities are postulated.

axioms (4)
  • standard math Tromborg-Waldenström bound on diagonal elements of a unitary matrix
    Used in Lemma 1 proof (Appendix A, Eq. 48) to bound |Tr(V†W)|.
  • domain assumption Loss is infidelity and optimization proceeds over all unitaries up to global phase
    Sections 2.2 and 4; defines the problem being analyzed.
  • domain assumption Reference system dimension equals main system dimension, pure training states
    Section 2.1; used to derive Eq. 23, the fidelity-distance relation.
  • domain assumption The local-neighborhood metric d'_F on PU(d) captures the geometry of PQC loss landscapes
    Sections 4 and 5; assumed for analytic results, validated only by n=5 simulations.

pith-pipeline@v1.3.0-alltime-deepseek · 38337 in / 13516 out tokens · 135112 ms · 2026-08-04T18:05:30.587310+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Loss Behavior in Supervised Learning with Entangled States." pith.science (2026). https://pith.science/paper/AMVW4GUQ

@misc{pith2026250910141,
  author       = {Pith},
  title        = {Pith review of: Loss Behavior in Supervised Learning with Entangled States},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AMVW4GUQ}},
  note         = {Machine review of arXiv:2509.10141}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Quantum Machine Learning (QML) aims to leverage the principles of quantum mechanics to speed up the process of solving machine learning problems or improve the quality of solutions. Among these principles, entanglement with an auxiliary system was shown to increase the quality of QML models in applications such as supervised learning. Recent works focus on the information that can be extracted from entangled training samples and their effect on the approximation error of the trained model. However, results on the trainability of QML models show that the training process itself is affected by various properties of the supervised learning task. These properties include the circuit structure of the QML model, the used cost function, and noise on the quantum computer. To evaluate the applicability of entanglement in supervised learning, we augment these results by investigating the effect of highly entangled training data on the model's trainability. In this work, we show that for highly expressive models, i.e., models capable of expressing a large number of candidate solutions, the possible improvement of loss function values in constrained neighborhoods during optimization is severely limited when maximally entangled states are employed for training. Furthermore, we support this finding experimentally by simulating training with Parameterized Quantum Circuits (PQCs). Our findings show that as the expressivity of the PQC increases, it becomes more susceptible to loss concentration induced by entangled training data. Lastly, our experiments evaluate the efficacy of non-maximal entanglement in the training samples and highlight the fundamental role of entanglement entropy as a predictor for the trainability.

Figures

Figures reproduced from arXiv: 2509.10141 by Alexander Mandl, Frank Leymann, Johanna Barzen, Lavinia Stiliadou, Marvin Bechtold.

Figure 1
Figure 1. Figure 1: Comparison of loss landscapes for separable (left) and maximally entangled (right) training samples. [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: This plot summarizes the analytical results in Section 4. It shows the minimal obtainable training loss [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: The experiment setup for evaluating the loss landscape for a training sample [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Minimal losses when training with a separable state (blue) and a maximally entangled state (red). By [PITH_FULL_IMAGE:figures/full_fig_p015_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Distance to the closest minimum after optimizing in [PITH_FULL_IMAGE:figures/full_fig_p016_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Improvement of the loss function in the ball [PITH_FULL_IMAGE:figures/full_fig_p017_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Improvement for NME states. Each box shows the evaluated improvement in a local neighborhood [PITH_FULL_IMAGE:figures/full_fig_p018_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Improvement for NME states from Figure 7 reordered based on the Schmidt rank of each training [PITH_FULL_IMAGE:figures/full_fig_p019_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Circuit diagrams of the PQCs evaluated in this work. Each contained rotation gate ( [PITH_FULL_IMAGE:figures/full_fig_p040_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Evaluated expressivity for CRX-entanglement with l = 8 layers depending on the number of fidelity￾samples N. The markers indicate the average expressivity value across five evaluations, and the bars show the standard deviation. of the fidelity of Haar random states with the distribution of fidelities of states obtained from randomly selecting PQC parameters. Formally, it is estimated as the Kullback-Leibl… view at source ↗
Figure 11
Figure 11. Figure 11: Expressivity of the circuits from Figure 9 used in Section 5. Lower values correspond to a higher [PITH_FULL_IMAGE:figures/full_fig_p042_11.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

67 extracted references · 13 canonical work pages · 2 internal anchors

  1. [1]

    Quantum Science and Technology

    Maria Schuld and Francesco Petruccione.Supervised Learning with Quantum Computers. Quantum Science and Technology. Springer International Publishing, 2018. DOI: 10.1007/978- 3-319-96424-9

  2. [2]

    Quantum machine learning.Nature, 549(7671):195–202, 2017

    Jacob Biamonte, Peter Wittek, Nicola Pancotti, Patrick Rebentrost, Nathan Wiebe, and Seth Lloyd. Quantum machine learning.Nature, 549(7671):195–202, 2017. DOI: 10.1038/na- ture23474

  3. [3]

    Information-theoretic bounds on quan- tum advantage in machine learning.Physical review letters, 126 19:190505, 2021

    Hsin-Yuan Huang, Richard Kueng, and John Preskill. Information-theoretic bounds on quan- tum advantage in machine learning.Physical review letters, 126 19:190505, 2021. DOI: 10.1103/PhysRevLett.126.190505

  4. [4]

    Quantum support vector machine for big data classification.Physical review letters, 113(13):130503, 2014

    Patrick Rebentrost, Masoud Mohseni, and Seth Lloyd. Quantum support vector machine for big data classification.Physical review letters, 113(13):130503, 2014. DOI: 10.1103/Phys- RevLett.113.130503

  5. [5]

    Classification with quantum neural networks on near term processors.arXiv:1802.06002, 2018

    Edward Farhi and Hartmut Neven. Classification with quantum neural networks on near term processors.arXiv:1802.06002, 2018. DOI: 10.48550/arXiv.1802.06002

  6. [6]

    A quantum approximate optimization algorithm.arXiv:1411.4028, 2014

    Edward Farhi, Jeffrey Goldstone, and Sam Gutmann. A quantum approximate optimization algorithm.arXiv:1411.4028, 2014. DOI: 10.48550/arXiv.1411.4028

  7. [7]

    Warm-starting quantum optimization

    Daniel J Egger, Jakub Mareček, and Stefan Woerner. Warm-starting quantum optimization. Quantum, 5:479, 2021. DOI: 10.22331/q-2021-06-17-479

  8. [8]

    Kyle Poland, Kerstin Beer, and Tobias J. Osborne. No free lunch for quantum machine learning.arXiv:2003.14103, 2020. DOI: 10.48550/arXiv.2003.14103

  9. [9]

    Cerezo, and Patrick J

    Kunal Sharma, Sumeet Khatri, M. Cerezo, and Patrick J. Coles. Noise resilience of varia- tional quantum compiling.New Journal of Physics, 22(4):043006, 2020. DOI: 10.1088/1367- 2630/ab784c

  10. [10]

    Sornborger, and Patrick J

    Sumeet Khatri, Ryan LaRose, Alexander Poremba, Lukasz Cincio, Andrew T. Sornborger, and Patrick J. Coles. Quantum-assisted quantum compiling.Quantum, 3:140, 2019. DOI: 10.22331/q-2019-05-13-140

  11. [11]

    Parameterized quantum circuits as machine learning models.Quantum Science and Technology, 4(4):043001, 2019

    Marcello Benedetti, Erika Lloyd, Stefan Sack, and Mattia Fiorentini. Parameterized quantum circuits as machine learning models.Quantum Science and Technology, 4(4):043001, 2019. DOI: 10.1088/2058-9565/ab4eb5

  12. [12]

    Quantum advantage in learning from experiments.Science, 376(6598):1182–1186, 2022

    Hsin-Yuan Huang, Michael Broughton, Jordan Cotler, et al. Quantum advantage in learning from experiments.Science, 376(6598):1182–1186, 2022. DOI: 10.1126/science.abn7293

  13. [13]

    Caro, Hsin-Yuan Huang, Nicholas Ezzell, et al

    Matthias C. Caro, Hsin-Yuan Huang, Nicholas Ezzell, et al. Out-of-distribution general- ization for learning quantum dynamics.Nature Communications, 14(1):3751, 2023. DOI: 10.1038/s41467-023-39381-w

  14. [14]

    Universal compiling and (no-)free-lunch theorems for continuous-variable quantum learning.PRX Quantum, 2:040327, 2021

    Tyler Volkoff, Zoë Holmes, and Andrew Sornborger. Universal compiling and (no-)free-lunch theorems for continuous-variable quantum learning.PRX Quantum, 2:040327, 2021. DOI: 10.1103/PRXQuantum.2.040327

  15. [15]

    Cerezo, Zoë Holmes, Lukasz Cincio, Andrew Sornborger, and Patrick J

    Kunal Sharma, M. Cerezo, Zoë Holmes, Lukasz Cincio, Andrew Sornborger, and Patrick J. Coles. Reformulation of the no-free-lunch theorem for entangled datasets.Physical Review Letters, 128(7), February 2022. DOI: 10.1103/physrevlett.128.070501

  16. [16]

    Transition role of entangled data in quantum machine learning.Nature Communications, 15(1):3716,

    Xinbiao Wang, Yuxuan Du, Zhuozhuo Tu, Yong Luo, Xiao Yuan, and Dacheng Tao. Transition role of entangled data in quantum machine learning.Nature Communications, 15(1):3716,

  17. [17]

    On Reducing the Amount of Samples Required for Training of QNNs: Constraints on the Linear Structure of the Training Data

    Alexander Mandl, Johanna Barzen, Frank Leymann, and Daniel Vietz. On reducing the amount of samples required for training of QNNs: Constraints on the linear structure of the training data.arXiv:2309.13711, 2023. DOI: 10.48550/arXiv.2309.13711

  18. [18]

    Cerezo, and Patrick J

    Zoë Holmes, Kunal Sharma, M. Cerezo, and Patrick J. Coles. Connecting ansatz express- ibility to gradient magnitudes and barren plateaus.PRX Quantum, 3(1):010313, 2022. DOI: 10.1103/PRXQuantum.3.010313

  19. [19]

    Barren plateaus in variational quantum computing.Nature Reviews Physics, 2025

    Martín Larocca, Supanut Thanasilp, Samson Wang, et al. Barren plateaus in variational quantum computing.Nature Reviews Physics, 2025. DOI: 10.1038/s42254-025-00813-9

  20. [20]

    Cerezo, Akira Sone, Tyler Volkoff, Lukasz Cincio, and Patrick J

    M. Cerezo, Akira Sone, Tyler Volkoff, Lukasz Cincio, and Patrick J. Coles. Cost function de- pendent barren plateaus in shallow parametrized quantum circuits.Nature Communications, 12(1):1791, 2021. DOI: 10.1038/s41467-021-21728-w. 23

  21. [21]

    Cerezo, et al

    Samson Wang, Enrico Fontana, M. Cerezo, et al. Noise-induced barren plateaus in variational quantum algorithms.Nature Communications, 12(1):6961, 2021. DOI: 10.1038/s41467-021- 27045-6

  22. [22]

    Lorenzo Leone, Salvatore FE Oliviero, Lukasz Cincio, and M. Cerezo. On the practical use- fulness of the hardware efficient ansatz.Quantum, 8:1395, 2024. DOI: 10.22331/q-2024-07- 03-1395

  23. [23]

    Supanut Thanasilp, Samson Wang, Nhat Anh Nghiem, Patrick Coles, and M. Cerezo. Sub- tleties in the trainability of quantum machine learning models.Quantum Machine Intelligence, 5(1), 2023. DOI: 10.1007/s42484-023-00103-6

  24. [24]

    Exploring the cost landscape of variational quantum algorithms

    Lavinia Stiliadou, Johanna Barzen, Frank Leymann, Alexander Mandl, and Benjamin Weder. Exploring the cost landscape of variational quantum algorithms. InSymposium and Summer School on Service-Oriented Computing, pages 128–142. Springer, 2024. DOI: 10.1007/978-3- 031-72578-4_7

  25. [25]

    Cerezo, and Patrick J

    Andrew Arrasmith, Zoë Holmes, M. Cerezo, and Patrick J. Coles. Equivalence of quantum barren plateaus to cost concentration and narrow gorges.Quantum Science and Technology, 7(4):045015, 2022. DOI: 10.1088/2058-9565/ac7d06

  26. [26]

    A unifying account of warm start guarantees for patches of quantum landscapes.arXiv:2502.07889, 2025

    Hela Mhiri, Ricard Puig, Sacha Lerch, Manuel S Rudolph, Thiparat Chotibut, Supanut Thanasilp, and Zoë Holmes. A unifying account of warm start guarantees for patches of quantum landscapes.arXiv:2502.07889, 2025. DOI: 10.48550/arXiv.2502.07889

  27. [27]

    Parallel variational quantum algorithms with gradient-informed restart to speed up optimisation in the presence of barren plateaus

    Daniel Mastropietro, Georgios Korpas, Vyacheslav Kungurtsev, and Jakub Marecek. Parallel variational quantum algorithms with gradient-informed restart to speed up optimisation in the presence of barren plateaus.arXiv:2311.18090, 2023. DOI: 10.48550/arXiv.2311.18090

  28. [28]

    Nielsen and Isaac L

    Michael A. Nielsen and Isaac L. Chuang.Quantum Computation and Quantum Information. Cambridge University Press, 2010. DOI: 10.1017/CBO9780511976667

  29. [29]

    Concen- trating partial entanglement by local operations.Physical Review A, 53(4):2046, 1996

    Charles H Bennett, Herbert J Bernstein, Sandu Popescu, and Benjamin Schumacher. Concen- trating partial entanglement by local operations.Physical Review A, 53(4):2046, 1996. DOI: 10.1103/PhysRevA.53.2046

  30. [30]

    CambridgeUniversityPress, 2017

    Ingemar Bengtsson and Karol Życzkowski.Geometry of quantum states: an introduction to quantum entanglement. CambridgeUniversityPress, 2017. DOI:10.1017/CBO9780511535048

  31. [31]

    Introduction to haar measure tools in quantum information: A beginner’s tutorial.Quantum, 8:1340, 2024

    Antonio Anna Mele. Introduction to haar measure tools in quantum information: A beginner’s tutorial.Quantum, 8:1340, 2024. DOI: 10.22331/q-2024-05-08-1340

  32. [32]

    Miszczak

    Zbigniew Puchała and Jarosław A. Miszczak. Symbolic integration with respect to the Haar measure on the unitary groups.Bulletin of the Polish Academy of Sciences Technical Sciences, 65(1):21–27, 2017. DOI: 10.1515/bpasts-2017-0003

  33. [33]

    Quantum speed limit based on the bound of Bures angle

    Shao-xiong Wu and Chang-shui Yu. Quantum speed limit based on the bound of Bures angle. Scientific reports, 10(1):5500, 2020. DOI: 10.1038/s41598-020-62409-w

  34. [34]

    Lie groups, Lie algebras, and representations

    Brian C Hall. Lie groups, Lie algebras, and representations. InQuantum Theory for Mathe- maticians, pages 333–366. Springer, 2013. DOI: 10.1007/978-3-319-13467-3

  35. [35]

    Query-optimal es- timation of unitary channels in diamond distance

    Jeongwan Haah, Robin Kothari, Ryan O’Donnell, and Ewin Tang. Query-optimal es- timation of unitary channels in diamond distance. In2023 IEEE 64th Annual Sympo- sium on Foundations of Computer Science (FOCS), pages 363–390. IEEE, 2023. DOI: 10.1109/FOCS57990.2023.00028

  36. [36]

    Johnson, and Alán Aspuru-Guzik

    Sukin Sim, Peter D. Johnson, and Alán Aspuru-Guzik. Expressibility and entangling capa- bility of parameterized quantum circuits for hybrid quantum-classical algorithms.Advanced Quantum Technologies, 2(12):1900070, 2019. DOI: 10.1002/qute.201900070

  37. [37]

    Cerezo et al

    M. Cerezo et al. Variational quantum algorithms.Nature Reviews Physics, 3(9):625–644,

  38. [38]

    On the differential topology of expressivity of parame- terized quantum circuits.AppliedMath, 5(3):121, 2025

    Johanna Barzen and Frank Leymann. On the differential topology of expressivity of parame- terized quantum circuits.AppliedMath, 5(3):121, 2025. DOI: 10.3390/appliedmath5030121

  39. [39]

    Quantum circuit ansatz: patterns of ab- straction and reuse of quantum algorithm design

    Xiaoyu Guo, Takahiro Muta, and Jianjun Zhao. Quantum circuit ansatz: patterns of ab- straction and reuse of quantum algorithm design. In2024 IEEE International Conference on Quantum Software (QSW), pages 69–80. IEEE, 2024. DOI: 10.1109/QSW62656.2024.00021

  40. [40]

    Variational quantum eigensolver with linear depth problem-inspired ansatz for solving portfolio optimization in finance.Science China Information Sciences, 68(8):1–11, 2025

    Shengbin Wang, Peng Wang, Guihui Li, et al. Variational quantum eigensolver with linear depth problem-inspired ansatz for solving portfolio optimization in finance.Science China Information Sciences, 68(8):1–11, 2025. DOI: 10.1007/s11432-024-4185-1. 24

  41. [41]

    Hardware-efficient variational quantum eigensolver for small molecules and quantum magnets.Nature, 549(7671):242–246,

    Abhinav Kandala, Antonio Mezzacapo, Kristan Temme, et al. Hardware-efficient variational quantum eigensolver for small molecules and quantum magnets.Nature, 549(7671):242–246,

  42. [42]

    Dimensional expressivity analysis of parametric quantum circuits.Quantum, 5:422, 2021

    Lena Funcke, Tobias Hartung, Karl Jansen, Stefan Kühn, and Paolo Stornati. Dimensional expressivity analysis of parametric quantum circuits.Quantum, 5:422, 2021

  43. [43]

    Quantum neural network cost function concentration dependency on the parametrization expressivity.Scientific reports, 13(1):9978, 2023

    Lucas Friedrich and Jonas Maziero. Quantum neural network cost function concentration dependency on the parametrization expressivity.Scientific reports, 13(1):9978, 2023. DOI: 10.1038/s41598-023-37003-5

  44. [44]

    Roughness index for loss landscapes of neural network models of partial differential equations

    Keke Wu, Xiangru Jian, Rui Du, Jingrun Chen, and Xiang Zhou. Roughness index for loss landscapes of neural network models of partial differential equations. In2023 IEEE Inter- national Conference on Big Data (BigData), pages 966–975. IEEE, 2023. DOI: 10.1109/Big- Data59044.2023.10386909

  45. [45]

    Cerezo, Samson Wang, Tyler Volkoff, Andrew T

    Arthur Pesah, M. Cerezo, Samson Wang, Tyler Volkoff, Andrew T. Sornborger, and Patrick J. Coles. Absence of barren plateaus in quantum convolutional neural networks.Physical Review X, 11(4), 2021. DOI: 10.1103/physrevx.11.041011

  46. [46]

    Barren plateaus in quantum neural network training landscapes.Nature Communications, 9 (1):4812, 2018

    Jarrod R McClean, Sergio Boixo, Vadim N Smelyanskiy, Ryan Babbush, and Hartmut Neven. Barren plateaus in quantum neural network training landscapes.Nature Communications, 9 (1):4812, 2018. DOI: 10.1038/s41467-018-07090-4

  47. [47]

    An initial- ization strategy for addressing barren plateaus in parametrized quantum circuits.Quantum, 3:214, 2019

    Edward Grant, Leonard Wossnig, Mateusz Ostaszewski, and Marcello Benedetti. An initial- ization strategy for addressing barren plateaus in parametrized quantum circuits.Quantum, 3:214, 2019. DOI: 10.22331/q-2019-12-09-214

  48. [48]

    Estimates of loss function concentra- tion in noisy parametrized quantum circuits.arXiv preprint arXiv:2410.01893, 2024

    Giulio Crognaletti, Michele Grossi, and Angelo Bassi. Estimates of loss function concentra- tion in noisy parametrized quantum circuits.arXiv preprint arXiv:2410.01893, 2024. DOI: 10.48550/arXiv.2410.01893

  49. [49]

    Data repository for ’Loss Behavior in Supervised Learning With Entangled States’

    Alexander Mandl, Johanna Barzen, Marvin Bechtold, Frank Leymann, and Lavinia Stiliadou. Data repository for ’Loss Behavior in Supervised Learning With Entangled States’. 2023. DOI: 10.18419/DARUS-5174

  50. [50]

    Pytorch: An imperative style, high- performance deep learning library

    Adam Paszke, Sam Gross, Francisco Massa, et al. Pytorch: An imperative style, high- performance deep learning library. InAdvances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019

  51. [51]

    A software package for sequential quadratic programming.Forschungsbericht- Deutsche Forschungs- und Versuchsanstalt fur Luft- und Raumfahrt, 1988

    Dieter Kraft. A software package for sequential quadratic programming.Forschungsbericht- Deutsche Forschungs- und Versuchsanstalt fur Luft- und Raumfahrt, 1988

  52. [52]

    SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python.Nature Methods, 17:261–272, 2020

    Pauli Virtanen, Ralf Gommers, Travis E Oliphant, et al. SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python.Nature Methods, 17:261–272, 2020. DOI: 10.1038/s41592- 019-0686-2

  53. [53]

    Elementary gates for quantum computation.Physical review A, 52(5):3457, 1995

    Adriano Barenco, Charles H Bennett, Richard Cleve, et al. Elementary gates for quantum computation.Physical review A, 52(5):3457, 1995. DOI: 10.1103/PhysRevA.52.3457

  54. [54]

    Warm-starting and quantum comput- ing: A systematic mapping study.ACM Comput

    Felix Truger, Johanna Barzen, Marvin Bechtold, et al. Warm-starting and quantum comput- ing: A systematic mapping study.ACM Comput. Surv., 56(9), 2024. DOI: 10.1145/3652510

  55. [55]

    Cerezo, Martin Larocca, Diego García-Martín, et al

    M. Cerezo, Martin Larocca, Diego García-Martín, et al. Does provable absence of barren plateaus imply classical simulability?Or, why we need to rethink variational quantum com- puting.arXiv preprint arXiv:2312.09121, 2023. DOI: 10.48550/arXiv.2312.09121

  56. [56]

    Bounds on the diagonal elements of a unitary matrix.Linear Algebra and its Applications, 20(3):189–195, 1978

    B Tromborg and S Waldenstrøm. Bounds on the diagonal elements of a unitary matrix.Linear Algebra and its Applications, 20(3):189–195, 1978. DOI: 10.1016/0024-3795(78)90017-4

  57. [57]

    Springer, 2015

    Jörg Liesen and Volker Mehrmann.Linear Algebra. Springer, 2015. DOI: 10.1007/978-3-319- 24346-7

  58. [58]

    I. M. Gelfand and Mark Saul.Trigonometry. Birkhäuser Boston, 2001. ISBN 9781461201496. DOI: 10.1007/978-1-4612-0149-6

  59. [59]

    Certain inequalities of kober and lazarevic type.The Journal of the Indian Mathematical Society, 89(1-2):01–07, 2022

    Yogesh J Bagul and Satish K Panchal. Certain inequalities of kober and lazarevic type.The Journal of the Indian Mathematical Society, 89(1-2):01–07, 2022. DOI: 10.18311/jims/2022/20737

  60. [60]

    Jordan’s inequality: refinements, generalizations, applications and related problems

    Feng Qi. Jordan’s inequality: refinements, generalizations, applications and related problems. Research report collection, 9(3), 2006

  61. [61]

    Expressibility of the alternating layered ansatz for quantum computation.Quantum, 5:434, 2021

    Kouhei Nakaji and Naoki Yamamoto. Expressibility of the alternating layered ansatz for quantum computation.Quantum, 5:434, 2021. DOI: 10.22331/q-2021-04-19-434. 25

  62. [62]

    Average fidelity between random quantum states.Phys

    Karol Życzkowski and Hans-Jürgen Sommers. Average fidelity between random quantum states.Phys. Rev. A, 71:032313, Mar 2005. DOI: 10.1103/PhysRevA.71.032313

  63. [63]

    Random pure states

    Sumeet Khatri. Random pure states. Notes available onlinehttps://sumeetkhatri.com/ notes/, 2020. Accessed 2025-09-09. A Distance to Fixed-Fidelity Operators In this section, we evaluate the Frobenius norm distanced′ F (V,W)from a starting pointV∈PU(d) toanyoperatorWwitha fixedtarget fidelityF U,ψ(W) =f W. Thus, wefind boundson thedistance to the closest o...

  64. [67]

    By rewriting the inner product in Equation (51) using these decompositions and applying the triangle inequality, we have ⏐⏐⟨ψ1|V†W|ψ 1⟩ ⏐⏐= ⏐⏐⟨ψ|V†UU†W|ψ⟩ ⏐⏐ (54) = ⏐⏐⏐ ( U†V|ψ⟩ )†( U†W|ψ⟩ )⏐⏐⏐ (55) = ⏐⏐⏐ ( ⟨ψ|U†V|ψ⟩|ψ⟩+|ψ ⊥ U †V⟩ )†( ⟨ψ|U†W|ψ⟩|ψ⟩+|ψ ⊥ U †W⟩ )⏐⏐⏐ (56) = ⏐⏐( ⟨ψ|V†U|ψ⟩⟨ψ|+⟨ψ ⊥ U †V| )( ⟨ψ|U†W|ψ⟩|ψ⟩+|ψ ⊥ U †W⟩ )⏐⏐ (57) = ⏐⏐⏐⟨ψ|U†W|ψ⟩⟨ψ|V †U|...

  65. [2017]

    DOI: 10.1038/nature23879

  66. [2021]

    DOI: 10.1038/s42254-021-00348-9

  67. [2024]

    DOI: 10.1038/s41467-024-47983-1