REVIEW 3 major objections 6 minor 1 cited by
Trainability of Parametrised Linear Combinations of Unitaries
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper proves that a linear combination of trainable parametrised circuits is itself trainable, with the variance of the combination lower-bounded by any component's variance up to a factor of $1/k^3$ in the number of terms.
desk verdict Solid variance formulas for the unnormalized LCU expectation, but the advertised trainability claim does not follow because the actual postselected cost is the normalized ratio. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the unnormalised expectation $m = tr(O|\psi\rangle\langle\psi|)$, whose subnormalisation encodes the LCU postselection probability. The argument expands $m = \sum_{i,j} c_i c_j m_{ij}$ with $m_{ij} = tr(U_i \rho_0 U_j^\dagger O)$, then expands the variance of $m$ into coefficient moments and unitary moments. Coefficient moments come from the uniform Dirichlet distribution; unitary moments come from Haar/Weingarten calculus on the Lie group generated by each $U_j$. For the matchgate case the relevant group is $SO(2N)$, whose commutant has the two-element basis identity and fermionic parity, which makes $E[|m_{ij}|^2]$ exactly computable and yields the closed-form variance formula.
What would settle it
Run a numerical Monte Carlo evaluation of the variance of the normalised postselected expectation $tr(O|\psi\rangle\langle\psi|)/\langle\psi|\psi\rangle$ for uniform Dirichlet LCUs of Haar-random matchgate unitaries at $N \ge 10$; if that variance decays exponentially in $N$ for fixed $k$ while each component is trainable under the paper's definition, the stated trainability guarantee would not apply to the normalised objective.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that the LCU construction inherits trainability from its parts. With output state $|\psi\rangle = \sum_{j=1}^{k} c_j U_j|0\rangle$ and expectation $m = tr(O|\psi\rangle\langle\psi|)$, the proof shows that the variance of $m$ is at least $1/k^3$ times the smallest component variance whenever each single-circuit expectation has zero mean. Thus a polynomial lower bound on any component variance is carried over to the combination, with only a polynomial penalty in $k$. Under uniform Dirichlet coefficients the variance is given exactly by a weighted sum of single-term moments, cross-term products, and inter-term moments $E[|m_{ij}|^2]$, the last computable by Weingarten calculus over the relevant Lie group. The same reasoning, with a stronger bound, applies to incoherent superpositions.
Load-bearing premise
The argument identifies trainability with the variance of the unnormalised expectation $tr(O|\psi\rangle\langle\psi|)$; if the quantity actually optimised in an LCU application is the normalised postselected expectation $tr(O|\psi\rangle\langle\psi|)/\langle\psi|\psi\rangle$, the paper's bound does not by itself constrain that ratio's variance.
Editorial extensions
If this is right
- If each component circuit has variance $\Omega(1/N^s)$, an LCU of $k$ such circuits has variance $\Omega(1/(N^s k^3))$, so the combined circuit is trainable in the polynomial sense.
- The exact Dirichlet formula means the variance can be computed analytically for any LCU whose single-term and cross-term moments are known, so trainability is certifiable term by term.
- For matchgate blocks, increasing the number of terms $k$ increases the fermionic Gaussian rank, a measure of expressivity, while the variance bound stays polynomial in $N$, yielding a tunable family of expressive, trainable circuits.
- Incoherent superpositions inherit trainability even more strongly: the bound is $\Omega(1/(N^s k))$ rather than $\Omega(1/(N^s k^3))$.
- Matchgate LCU circuits cost $O(k N^2)$ quantum gates versus $O(k^2 N^3)$ classical operations for rank-$k$ fermionic Gaussian simulation, leaving an asymptotic quantum speedup, although on current hardware the classical gate-time advantage makes the speedup impractical.
Reading between the lines
- The paper measures trainability through the unnormalised expectation; a natural next step, not taken here, is to prove or test the same lower bound for the normalised postselected estimator, which is what an LCU-based optimiser actually evaluates.
- The worst-case $1/k^3$ factor implies a direct expressivity-trainability trade-off: increasing the number of terms $k$ enriches the state family, but for fixed component variances it shrinks the guaranteed variance, so ansatz design would tune $k$ against the components' variance decay.
- One could apply the bound recursively to hierarchical LCUs, linear combinations whose components are themselves linear combinations, since the theorem only requires each component to have a polynomial variance, not any particular internal structure.
- Sampling coefficients from a non-uniform Dirichlet distribution that concentrates weight on the more trainable components might improve the effective constant beyond the uniform worst case; this is a direct, testable modification of the paper's setup.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the trainability of parametrised linear combinations of unitaries (LCUs). The setting is a state |ψ> = Σ_j c_j U_j|0> with l1-normalised coefficients c and Haar-random unitaries U_j over a Lie group, and the quantity of interest is m = tr(O|ψ><ψ|) with |ψ> left subnormalised. The main results are Theorem 1, a lower bound Var[m] ≥ (1/k^3) min_j Var[m_j] independent of the coefficient distribution, and Theorem 2, an exact expression for Var[m] when c is uniform Dirichlet. These are specialised to fermionic Gaussian unitaries (matchgates), giving Corollaries 1 and 2 which state a Ω(1/(N^s k^3)) scaling, and to Haar-random unitaries. Numerical experiments on matchgate LCUs are reported as agreeing with the analytical formula. The paper also extends the bound and exact formula to incoherent superpositions.
Significance. If the central claim were established, the paper would be a useful contribution: it would extend barren-plateau analysis to a widely used construction, provide a family of trainable circuits with tunable expressivity, and give concrete analytical variance formulas that are checked numerically. The algebraic machinery (Dirichlet moments, Weingarten computation for SO(2N)) is largely standard, and the numerical check against the analytical formula is a genuine strength. However, the advertised conclusion that an 'LCU of trainable circuits is still trainable' for the postselected LCU cost is not supported by the variance analysis of the unnormalised expectation, and the claimed gradient lower bound is not proved. The exact variance formulas also appear to omit a term that does not vanish in general, so the paper needs substantial revision before its central claims can be accepted.
major comments (3)
- [Section 2, Eq. (3) and Theorem 1 (Eq. (7))] The paper defines m = tr(O|ψ><ψ|) with |ψ> = Σ_j c_j U_j|0> subnormalised, and states that not normalising the state 'takes into account the postselection probability'. In actual LCU applications the cost is the postselected normalised expectation M = tr(O|ψ><ψ|)/<ψ|ψ>, where q = <ψ|ψ> is the success probability. The variance of m does not control the variance of M: the denominator q is a random variable (under uniform Dirichlet coefficients E[q] = 2/(k+1)), and M can have large variance even when Var[m] is polynomially small, because q can be close to zero. Theorems 1 and 2 bound only Var[m]. No bound, expression, or analysis is given for Var[M]. Consequently, the paper does not establish that an LCU of trainable circuits is trainable for the cost actually evaluated in an LCU protocol. The revision should either restrict all claims to the unnormalised cost, clearly labelling it as such, or supply a rigorous variance analysis of M and its gradient.
- [Section 3.1, final paragraph] The text concludes: 'We conclude that, for linear combinations of fermionic Gaussian states, the gradient has a lower bound polynomial in both the number of qubits, and the rank of the system.' This does not follow from Theorem 2 or Corollary 1, which bound Var[m] for the expectation value, not Var[∂_θ m] for any parameter derivative. In the barren-plateau literature trainability is defined through the gradient variance; a bound on the cost variance does not by itself imply a bound on the gradient variance. The revision should either provide a proof connecting the two (for example, via a parameter-shift rule and a triangle inequality applied to the unnormalised or normalised cost) or remove this claim.
- [Appendix A.1, Lemma 3 and Eq. (37)] The proof of Lemma 3 uses invariance of the Haar measure under left multiplication by e^{iπ/n} I. This element is not contained in SO(2N) for n=2, and it is not contained in SU(2) for n=2, so the conclusion E[m_ij^n]=0 for n≥2 does not follow for the matchgate case or for N=1 unitary case. As a result, the reduction of E[m^2] to the three sums in Eq. (37) omits the term Σ_{i≠j} E[c_i^2 c_j^2] E[m_ij^2]. This term is nonzero already for SU(2): taking ρ0=I/2, O=Z, k=2, and deterministic c=(1/2,1/2), direct computation gives m = 0 deterministically, so Var[m]=0, whereas Eq. (9) gives 1/60. Thus Theorem 2 and Corollaries 1 and 3 are incomplete as stated. Please either prove that E[m_ij^2]=0 for the groups considered (with a correct argument) or include the missing term in the exact expressions and re-derive the corollaries.
minor comments (6)
- [Section 1, final paragraph] There is a typo: 'postselction' should be 'postselection'.
- [Appendix E.1, Eq. (118)] The second term inside the brackets uses 2^{N-2}, whereas the main-text Eq. (11) and the derivation from Eq. (51) give 2^{N-1}; this appears to be a typo in the appendix.
- [Theorem 2 statement] The phrase 'For state a ρ prepared' should read 'For a state ρ prepared'.
- [Appendix D, Eq. (105) and the paragraph before it] The symbol P is used for the fermionic parity operator in Lemma 1 and Corollary 1 before it is explicitly defined; please define P at first use.
- [Section 4, numerical results] The numerical agreement is visual only; no error bars or statistical error estimates are given. Please state the standard error of the mean or include error bars so that the agreement with the analytical formula can be assessed quantitatively.
- [Bibliography] The citation 'Diaz et al. [2023]' is used for two different references: 'Showcasing a barren plateau theory beyond the dynamical lie algebra' and 'Classical simulation of non-gaussian fermionic circuits' (Dias and König). Please disambiguate the names to avoid confusion.
Circularity Check
No significant circularity: the variance bounds are genuine analytic derivations from explicit Haar-measure and Dirichlet assumptions, with numerical checks rather than fitted predictions.
full rationale
The paper's derivation chain is self-contained: Theorem 1 follows from expanding Var[m] in terms of coefficient moments and E[m_i^2], then applying Holder's inequality (Lemma 4); Theorem 2 substitutes closed-form moments of the uniform Dirichlet distribution into that expansion. The SO(2N) specialization uses external moment formulas from Diaz et al. (2023) and standard Weingarten/Haar identities, and the numerics are consistency checks, not fits to the derived expression. The only self-citation (Khatri et al. 2024) appears as a motivating example and carries no load. The reviewer's main worry, that the bounded quantity is the unnormalized expectation m = qM rather than the actual postselected cost M = tr(O|psi><psi|)/<psi|psi>, is a substantive applicability gap in the advertised claim, but it is not a circularity: the theorems do not assume the conclusion about M, and the variance of m is derived, not defined, to be the trainability measure. A revision should either restrict the trainability claim to m or supply a bound on Var[M]; this does not make the existing derivation circular. Similarly, the remark that the gradient has a polynomial lower bound is an unsupported extra inference, not a derivation from the proved variance, and so it is a correctness gap rather than a circular step. No step reduces Eq. X to Eq. Y by construction, and no fitted parameter is relabeled as a prediction.
Assumptions & free parameters
assumptions (3)
- domain assumption Each unitary U_j is sampled independently from the Haar measure over the Lie group G_j generated by the parametrised circuit.
- domain assumption The LCU coefficients c follow a uniform Dirichlet distribution Dir(1,...,1), and are positive with l1 norm equal to 1.
- domain assumption For the lower bound in Theorem 1, each component expectation satisfies E[m_i] = 0.
Cite this review
Pith. "Pith review of Trainability of Parametrised Linear Combinations of Unitaries." pith.science (2026). https://pith.science/paper/A3JGF5RF
@misc{pith2026250622310,
author = {Pith},
title = {Pith review of: Trainability of Parametrised Linear Combinations of Unitaries},
year = {2026},
howpublished = {\url{https://pith.science/paper/A3JGF5RF}},
note = {Machine review of arXiv:2506.22310}
}
read the original abstract
A principal concern in the optimisation of parametrised quantum circuits is the presence of barren plateaus, which present fundamental challenges to the scalability of applications, such as variational algorithms and quantum machine learning models. Recent proposals for these methods have increasingly used the linear combination of unitaries (LCU) procedure as a core component. In this work, we prove that an LCU of trainable parametrised circuits is still trainable. We do so by analytically deriving the expression for the variance of the expectation when applying the LCU to a set of parametrised circuits, taking into account the postselection probability. These results extend to incoherent superpositions. We support our conclusions with numerical results on linear combinations of fermionic Gaussian unitaries (matchgate circuits). Our work shows that sums of trainable parametrised circuits are still trainable, and thus provides a method to construct new families of more expressive trainable circuits. We argue that there is a scope for a quantum speed-up when evaluating these trainable circuits on a quantum device.
Figures
Forward citations
Cited by 1 Pith paper
-
Stacking the Deck: Tunable Trainability in Stacked LCUs
Stacked LCUs of fermionic Gaussian unitaries give variance Ω(1/(n k^{3l})) against classical simulation O(k^{2l} n^3) and quantum gate count O(l k n^2), with layers l as the single dial.
Reference graph
Works this paper leans on
-
[1]
Berry, Nathan Wiebe, Jarrod McClean, Alexandru Paler, Austin Fowler, and Hartmut Neven
Ryan Babbush, Craig Gidney, Dominic W. Berry, Nathan Wiebe, Jarrod McClean, Alexandru Paler, Austin Fowler, and Hartmut Neven. Encoding electronic spectra in quantum circuits with linear t complexity. Phys. Rev. X, 8: 0 041015, Oct 2018. doi:10.1103/PhysRevX.8.041015. URL https://link.aps.org/doi/10.1103/PhysRevX.8.041015
-
[2]
J. Bardeen, L. N. Cooper, and J. R. Schrieffer. Microscopic theory of superconductivity. Physical Review, 106 0 (1): 0 162--164, April 1957. doi:10.1103/physrev.106.162. URL https://doi.org/10.1103/physrev.106.162
-
[3]
JAX : composable transformations of P ython+ N um P y programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake Vander P las, Skye Wanderman- M ilne, and Qiao Zhang. JAX : composable transformations of P ython+ N um P y programs, 2018. URL http://github.com/jax-ml/jax
2018
-
[4]
Olson, Matthias Degroote, Peter D
Yudong Cao, Jonathan Romero, Jonathan P. Olson, Matthias Degroote, Peter D. Johnson, Mária Kieferová, Ian D. Kivlichan, Tim Menke, Borja Peropadre, Nicolas P. D. Sawaya, Sukin Sim, Libor Veis, and Alán Aspuru-Guzik. Quantum chemistry in the age of quantum computing. Chemical Reviews, 119 0 (19): 0 10856–10915, August 2019. ISSN 1520-6890. doi:10.1021/acs....
-
[5]
Cerezo, Martin Larocca, Diego García-Martín, N
M. Cerezo, Martin Larocca, Diego García-Martín, N. L. Diaz, Paolo Braccia, Enrico Fontana, Manuel S. Rudolph, Pablo Bermejo, Aroosa Ijaz, Supanut Thanasilp, Eric R. Anschuetz, and Zoë Holmes. Does provable absence of barren plateaus imply classical simulability? or, why we need to rethink variational quantum computing, 2024. URL https://arxiv.org/abs/2312.09121
arXiv 2024
-
[6]
Andew M. Childs and Nathan Wiebe. H amiltonian S imulation U sing L inear C ombinations of U nitary O perations. Quantum Information and Computation, 12 0 (11 & 12), November 2012. ISSN 1533-7146. doi:10.26421/qic12.11-12. URL http://dx.doi.org/10.26421/QIC12.11-12
-
[7]
On the sample complexity of quantum boltzmann machine learning
Luuk Coopmans and Marcello Benedetti. On the sample complexity of quantum boltzmann machine learning. Communications Physics, 7 0 (1): 0 274, 2024
work page 2024
-
[8]
Gaussian decomposition of magic states for matchgate computations, 2024
Joshua Cudby and Sergii Strelchuk. Gaussian decomposition of magic states for matchgate computations, 2024. URL https://arxiv.org/abs/2307.12654
arXiv 2024
Show all 24 references
-
[9]
Classical simulation of non-gaussian fermionic circuits
Beatriz Dias and Robert K\"onig. Classical simulation of non-gaussian fermionic circuits. Quantum, 8: 0 1350, 2024
2024
-
[10]
N. L. Diaz, Diego García-Martín, Sujay Kazi, Martin Larocca, and M. Cerezo. Showcasing a barren plateau theory beyond the dynamical lie algebra, 2023. URL https://arxiv.org/abs/2310.11505
2023 arXiv
-
[11]
Introduction to Quantum Control and Dynamics
Domenico D’Alessandro. Introduction to Quantum Control and Dynamics. Chapman and Hall/CRC, July 2021. ISBN 9781003051268. doi:10.1201/9781003051268. URL http://dx.doi.org/10.1201/9781003051268
2021 doi
-
[12]
Characterizing barren plateaus in quantum ans\" a tze with the adjoint representation
Enrico Fontana, Dylan Herman, Shouvanik Chakrabarti, Niraj Kumar, Romina Yalovetzky, Jamie Heredge, Shree Hari Sureshbabu, and Marco Pistoia. Characterizing barren plateaus in quantum ans\" a tze with the adjoint representation. Nature Communications, 15 0 (1), August 2024. IS...
2024 doi
-
[13]
Nonunitary quantum machine learning
Jamie Heredge, Maxwell West, Lloyd Hollenberg, and Martin Sevior. Nonunitary quantum machine learning. Phys. Rev. Appl., 23: 0 044046, Apr 2025. doi:10.1103/PhysRevApplied.23.044046. URL https://link.aps.org/doi/10.1103/PhysRevApplied.23.044046
2025 doi
-
[14]
Sung, Kostyantyn Kechedzhi, Vadim N
Zhang Jiang, Kevin J. Sung, Kostyantyn Kechedzhi, Vadim N. Smelyanskiy, and Sergio Boixo. Quantum algorithms to simulate many-body physics of correlated fermions. Phys. Rev. Appl., 9: 0 044036, Apr 2018. doi:10.1103/PhysRevApplied.9.044036. URL https://link.aps.org/doi/10.1103...
2018 doi
-
[15]
0.167em O
R. 0.167em O. Jones. Density functional theory: Its origins, rise to prominence, and future. Reviews of Modern Physics, 87 0 (3): 0 897--923, August 2015. doi:10.1103/revmodphys.87.897. URL https://doi.org/10.1103/revmodphys.87.897
2015 doi
-
[16]
Quixer: A quantum transformer model
Nikhil Khatri, Gabriel Matos, Luuk Coopmans, and Stephen Clark. Quixer: A quantum transformer model. arXiv preprint arXiv:2406.04305, 2024
2024 arXiv
-
[17]
Efekan K\"okc\"u, Daan Camps, Lindsay Bassman Oftelie, J. K. Freericks, Wibe A. de Jong, Roel Van Beeumen, and Alexander F. Kemper. Algebraic compression of quantum circuits for hamiltonian evolution. Phys. Rev. A, 105: 0 032420, Mar 2022. doi:10.1103/PhysRevA.105.032420. URL ...
2022 doi
-
[18]
Coles, Lukasz Cincio, Jarrod R
Martín Larocca, Supanut Thanasilp, Samson Wang, Kunal Sharma, Jacob Biamonte, Patrick J. Coles, Lukasz Cincio, Jarrod R. McClean, Zoë Holmes, and M. Cerezo. Barren plateaus in variational quantum computing. Nature Reviews Physics, 7 0 (4): 0 174–189, March 2025. ISSN 2522-5820...
2025 doi
-
[19]
Barren plateaus in quantum neural network training landscapes
Jarrod R McClean, Sergio Boixo, Vadim N Smelyanskiy, Ryan Babbush, and Hartmut Neven. Barren plateaus in quantum neural network training landscapes. Nature communications, 9 0 (1): 0 4812, 2018
2018
-
[20]
Introduction to haar measure tools in quantum information: A beginner's tutorial
Antonio Anna Mele. Introduction to haar measure tools in quantum information: A beginner's tutorial. Quantum, 8: 0 1340, 2024
2024
-
[21]
Bakalov, Frédéric Sauvage, Alexander F
Michael Ragone, Bojko N. Bakalov, Frédéric Sauvage, Alexander F. Kemper, Carlos Ortiz Marrero, Martín Larocca, and M. Cerezo. A lie algebraic theory of barren plateaus for deep parameterized quantum circuits. Nature Communications, 15 0 (1), August 2024. ISSN 2041-1723. doi:10...
2024 doi
-
[22]
Improved simulation of quantum circuits dominated by free fermionic operations
Oliver Reardon-Smith, Micha Oszmaniec, and Kamil Korzekwa. Improved simulation of quantum circuits dominated by free fermionic operations. Quantum, 8: 0 1549, 2024
2024
-
[23]
Terhal and David P
Barbara M. Terhal and David P. DiVincenzo. Classical simulation of noninteracting-fermion quantum circuits. Phys. Rev. A, 65: 0 032325, Mar 2002. doi:10.1103/PhysRevA.65.032325. URL https://link.aps.org/doi/10.1103/PhysRevA.65.032325
2002 doi
-
[24]
Leslie G. Valiant. Quantum computers that can be simulated classically in polynomial time. In Proceedings of the Thirty-Third Annual ACM Symposium on Theory of Computing, STOC '01, page 114–123, New York, NY, USA, 2001. Association for Computing Machinery. ISBN 1581133499. doi...
2001
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.