REVIEW 2 major objections 3 minor 48 references
Challenges in Barren Plateau Mitigation with Dynamic Parameterized Quantum Circuits
T0 review · 2 major / 3 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read Dynamic parameterized quantum circuits cannot generically rescue variational algorithms from barren plateaus: faithful variants inherit the base circuit's trainability failure.
desk verdict The paper's best-supported point—cost anti-concentration doesn't imply trainability—holds up, but Lemma 1 has a real gap: faithfulness is only imposed for small σ, while the conclusion bounds the variance over all σ. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is purification combined with Pauli-path analysis. Each parameterized CPTP gadget is dilated to a unitary V(σ) acting on system plus ancillas, with the ancillas discarded at the end; a 2-design is a unitary distribution matching the first two Haar moments, which scrambles observables enough to induce barren plateaus. In the purified Heisenberg picture, the observable breaks into terms supported only on ancillas, whose coefficients depend only on the gadget parameters σ and are untouched by θ, and terms with system support that are scrambled by θ. The ancilla-supported terms supply observable cost variance but no θ-gradient variance, which is exactly why anti-concen
What would settle it
Take the paper's DC-HEA or DC-QAOA setup, sample (θ, σ) uniformly over the full parameter range—including σ far outside the faithful ball—and compute Varθ,σ[∂θμ C] via the parameter-shift rule. If any θ-gradient variance is polynomially large rather than O(b^{-n}), Lemma 1's conclusion fails; if it is exponentially small even under full-range σ sampling, the unstated support restriction is harmless.
Extended reading notes
Core claim
The central claim is that dynamic parameterized quantum circuits (DPQCs) provide no generic loophole around barren plateaus. A DPQC that is faithful to a base unitary ansatz inherits every one of the base circuit's untrainable θ parameters (Lemma 1). If only O(1) gadget layers are inserted and a long intervening unitary stretch forms a 2-design, all parameters in that stretch become untrainable, leaving Ω(nL) parameters stranded (Lemma 2). The mechanism is exposed by purifying each non-unitary gadget: the observable evolves into Pauli terms supported only on the discarded ancillas, with coefficients depending only on the gadget parameters, so the cost can anti-concentrate while θ-gradients r
Load-bearing premise
Lemma 1 assumes that O(c^{-n}) closeness of costs for small gadget parameters forces every θ-gradient variance to be exponentially small; this holds only if the gadget parameters are effectively supported inside that small faithful ball, a restriction the paper states nowhere and its own numerics do not enforce.
Editorial extensions
If this is right
- Proposals that append a single dynamic gadget layer (or O(1) layers) to a deep random or hardware-efficient ansatz cannot prevent most parameters from becoming untrainable; the gadget only rescues parameters that come after it, and only for the last O(log n) layers.
- Cost anti-concentration, measured by a polynomially lower-bounded variance, should not be used as evidence of trainability; a DPQC can pass that test while every θ-gradient stays exponentially suppressed.
- To have any hope of a trainable DPQC, the circuit must sacrifice faithfulness and insert non-unitary layers densely, so that no long contiguous unitary sub-array forms a 2-design.
- Because a purified DPQC is itself a unitary circuit on a larger register, any provably BP-free DPQC would implicitly yield a BP-free unitary construction; so DPQC mitigation is at least as hard as unitary ansatz design.
- Optimization dynamics will show no improvement from the added dynamic layer: in the paper's VQE and QAOA experiments, joint optimization over (θ, σ) did not beat the unitary baseline, and removing the layer changed the result only slightly.
Reading between the lines
- Flagged limitation, not the paper's claim: the proof of Lemma 1 itself notes that second- or higher-order σ derivatives could cancel its first-order bound; the paper sets that aside, so the strongest lemma rests on an acknowledged gap.
- Left implicit: Lemma 1's faithfulness bound only constrains small σ, while the numerical studies sample σ over the full range; if large-σ regions carry genuine θ-gradients, a non-faithful DPQC could still be trainable, so the 'at least as hard' conclusion may admit problem-specific exceptions.
- Neighboring connection: the ancilla-supported Pauli mechanism resembles the reason noise-induced shallow circuits become classically simulable; if the analogy holds, DPQCs that anti-concentrate for this reason may also be classically simulable in their trainable directions.
- Testable extension: use the paper's symbolic Pauli-Heisenberg engine to search gadget families for high Varσ[F(σ)] at low depth; the framework sets up the design tool but leaves systematic search, and the question of non-classicality in the σ directions, open.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper formalizes dynamic parameterized quantum circuits (DPQCs), i.e., unitary parameterized circuits interleaved with parameterized CPTP gadgets, and asks whether such circuits can mitigate barren plateaus (BPs). The main theoretical claims are: (Lemma 1) a DPQC that is faithful to a BP-prone unitary in the small-\sigma regime inherits the untrainability of the unitary's \theta parameters; (Lemma 2) inserting only O(1) gadget layers leaves all parameters in long 2-design unitary stretches untrainable; and (Lemma 3) cost anti-concentration can be certified from a classically computable ancilla-supported Pauli term F(\sigma). Numerical studies on VQE with a hardware-efficient ansatz and Max-Cut QAOA show anti-concentrated costs coexisting with exponentially small \theta-gradient variance. The authors conclude that BP mitigation via DPQCs is at least as hard as designing BP-free unitary ansatze.
Significance. If the results held, the paper would provide a useful unifying framework for non-unitary BP-mitigation proposals and would sharpen the distinction between cost anti-concentration and actual parameter trainability. The Pauli-path criterion in Lemma 3 is a valuable diagnostic, and the numerical work is carefully structured: the analytic \Theta(1/n) variance bound in Appendix A matches Fig. 3a, full-range \sigma sampling is used, and code plus a symbolic Pauli engine (sympauli) are provided. However, the paper's advertised central claim rests on Lemma 1, which is false as stated. The remaining results are narrower than the abstract's conclusion and do not, by themselves, establish that BP mitigation via DPQCs is as hard as designing BP-free unitaries.
major comments (2)
- [Sec. II.C, Eq. (9), Appendix A] Lemma 1 is not proven and is false as stated. The faithfulness condition (Eq. 9) bounds |C(\theta,\sigma)-C(\theta,0)| only for ||\sigma||<\epsilon~1/poly(n), and the Appendix A proof uses parameter shifts to obtain a pointwise bound on |\partial_{\theta_\mu}C(\theta,\sigma)-\partial_{\theta_\mu}C(\theta,0)| in the same small-\sigma region. But Lemma 1 concludes Var_{\theta,\sigma}[\partial_{\theta_\mu}C]\in O(b^{-n}) over the full joint distribution. No argument controls the region ||\sigma||\ge\epsilon, and the paper never specifies the \sigma distribution. The gap is not merely technical: take E(\sigma)=\lambda(\sigma)Reset+(1-\lambda(\sigma))Id with \lambda=0 on ||\sigma||<\epsilon and \lambda=1 for ||\sigma||>2\epsilon. This satisfies Eq. (9), yet for large \sigma the system is reset before the later unitary layers, making downstream \theta parameters trainable. A correct version re
- [Sec. III and Abstract] The high-level conclusion that 'BP mitigation via DPQCs is at least as hard as designing BP-free unitaries' is not supported once Lemma 1 is removed or restricted. Lemma 2 applies only when long unitary stretches form 2-designs, and Lemma 3 is a design criterion rather than an impossibility statement. The counterexample to Lemma 1 shows that a DPQC faithful only in a small-\sigma neighborhood can, in principle, restore trainability of downstream parameters. The abstract and discussion should be recalibrated to the actually proven statements.
minor comments (3)
- [Sec. II.B, Eq. (6)] The symbol U is used for both the original unitary and the DPQC channel U(\theta,\sigma). This overloading is confusing; a different symbol for the channel would clarify the composition in Eq. (6) and the purified form in Eq. (11).
- [Sec. II.B, Eq. (9)] The domain of \sigma and the probability measure over (\theta,\sigma) are never defined. Since Eq. (14) and the numerical experiments average over \sigma, the relationship between the small-\epsilon faithfulness ball and the sampling distribution should be stated explicitly.
- [Appendix A, Lemma 1 proof] The passage 'To first order, we can even bound the \sigma gradients' followed by the caveat about higher-order cancellations is an acknowledgment that the argument is incomplete. If Lemma 1 is repaired, this part should either be removed or made into a rigorous statement, since it does not support the lemma as written.
Circularity Check
No significant circularity: Lemmas 1–3 are derived from stated definitions and cited external theorems; the only self-citation is non-load-bearing.
full rationale
The paper's central derivation chain is self-contained rather than circular. Lemma 1 uses the faithfulness definition (Eq. 9) together with parameter-shift identities to bound derivative differences; this is a derived consequence, not a restatement of the definition, because faithfulness is defined through cost closeness while the conclusion concerns gradient variance. Lemma 2 is a 2-design variance calculation using external facts about random circuits and HEA, and the QAOA Lie-algebra claim is supported by three references, only one of which ([32]) is by the present authors; that self-citation is not load-bearing. Lemma 3 is a law-of-total-variance argument with an explicit 1-design condition, and the anti-concentration lower bounds in the numerics are computed from the explicit gadget coefficients (e.g., Eq. A17 and Var[F] ~ Θ(1/n)), not fitted to the simulation data. No parameter is fitted and then renamed as a prediction. The one notable weakness, the Lemma 1 proof gap concerning the contribution of large-σ regions to Var_{θ,σ}, is a rigor/correctness concern rather than a circular construction: the conclusion is not equivalent to the faithfulness premise by definition. The self-citation [32] therefore does not raise the circularity score beyond the minor range.
Assumptions & free parameters
free parameters (1)
- Faithfulness radius ε =
Θ(1/poly(n))
assumptions (5)
- domain assumption The generators P_{j,k} in each unitary block are local and square to the identity (Eq. 3)
- domain assumption Any sufficiently long contiguous sub-array of U(θ) forms an approximate 2-design (Lemma 2)
- domain assumption U(θ) forms a 1-design and V(σ)† H V(σ) has polynomially many Pauli terms (Lemma 3, Eq. 12)
- standard math Stinespring dilation realizes arbitrary CPTP maps with O(1) ancillas and discarding them recovers the channel (Eq. 11)
- ad hoc to paper Global σ-variance is controlled by the faithful small-σ region (implicit in Lemma 1)
Cite this review
Pith. "Pith review of Challenges in Barren Plateau Mitigation with Dynamic Parameterized Quantum Circuits." pith.science (2026). https://pith.science/paper/DEXF2CWK
@misc{pith2026260623751,
author = {Pith},
title = {Pith review of: Challenges in Barren Plateau Mitigation with Dynamic Parameterized Quantum Circuits},
year = {2026},
howpublished = {\url{https://pith.science/paper/DEXF2CWK}},
note = {Machine review of arXiv:2606.23751}
}
read the original abstract
Variational quantum algorithms (VQAs) are a promising paradigm for quantum advantage, yet their trainability is severely hampered by barren plateaus (BPs). Several recent works have proposed dynamic parameterized quantum circuits (DPQCs), which interleave unitary layers with parameterized CPTP maps, such as engineered dissipation, feedforward gadgets, and periodic resets, as a possible strategy for mitigating BPs. We unify this class of circuits into a formalization for DPQCs.We identify constraints on the nature and the structure of DPQCs if they are to prevent a significant number of parameters from becoming untrainable. Using purification and Pauli-path analysis, we further identify a mechanism by which the cost function can remain anti-concentrated even when many parameters remain untrainable. Our analysis reveals ways to design DPQCs that do not have an exponentially concentrated cost function, and our results suggest that BP mitigation via DPQCs is at least as hard as designing BP-free unitaries.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Cerezo, A
M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio,et al., Variational quantum algorithms, Nature Reviews Physics3, 625 (2021)
2021
-
[2]
sympauli
(a purified DPQCisa unitary circuit) and it is only strengthened by our analysis. To achieve genuine trainability ofθin a DPQC, one must sacrifice faith- fulness and insert gadget layers frequently enough that the augmented landscape avoids 2-design behavior from contiguous sub-circuits. The most promising direction, in our view, is to study DPQCs whose g...
2026
-
[3]
Peruzzo, J
A. Peruzzo, J. McClean, P. Shadbolt, M.-H. Yung, X.-Q. Zhou, P. J. Love, A. Aspuru-Guzik, and J. L. O’brien, A variational eigenvalue solver on a photonic quantum processor, Nature communications5, 4213 (2014)
2014
- [4]
-
[5]
J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Bab- bush, and H. Neven, Barren plateaus in quantum neural network training landscapes, Nature communications9, 4812 (2018)
2018
-
[6]
Tilly, H
J. Tilly, H. Chen, S. Cao, D. Picozzi, K. Setia, Y. Li, E. Grant, L. Wossnig, I. Rungger, G. H. Booth,et al., The variational quantum eigensolver: a review of meth- ods and best practices, Physics Reports986, 1 (2022)
2022
-
[7]
Ragone, B
M. Ragone, B. N. Bakalov, F. Sauvage, A. F. Kemper, C. Ortiz Marrero, M. Larocca, and M. Cerezo, A lie al- gebraic theory of barren plateaus for deep parameter- ized quantum circuits, Nature Communications15, 7172 (2024)
2024
-
[8]
Larocca, S
M. Larocca, S. Thanasilp, S. Wang, K. Sharma, J. Bia- monte, P. J. Coles, L. Cincio, J. R. McClean, Z. Holmes, and M. Cerezo, Barren plateaus in variational quantum computing, Nature Reviews Physics7, 174 (2025)
2025
Show all 48 references
-
[9]
S. Wang, E. Fontana, M. Cerezo, K. Sharma, A. Sone, L. Cincio, and P. J. Coles, Noise-induced barren plateaus in variational quantum algorithms, Nature communica- tions12, 6961 (2021)
2021
-
[10]
Cerezo, A
M. Cerezo, A. Sone, T. Volkoff, L. Cincio, and P. J. Coles, Cost function dependent barren plateaus in shal- low parametrized quantum circuits, Nature communica- tions12, 1791 (2021)
2021
-
[11]
Deshpande, M
A. Deshpande, M. Hinsche, K. Najafi, K. Sharma, R. Sweke, and C. Zoufal, Dynamic parameterized quan- tum circuits: expressive and barren-plateau free, arXiv preprint arXiv:2411.05760 (2024)
2024 arXiv
-
[12]
Arrasmith, Z
A. Arrasmith, Z. Holmes, M. Cerezo, and P. J. Coles, Equivalence of quantum barren plateaus to cost concen- tration and narrow gorges, Quantum Science & Technol- ogy7, 045015 (2022)
2022
-
[13]
Zhang, L
K. Zhang, L. Liu, M.-H. Hsieh, and D. Tao, Escaping from the barren plateau via gaussian initializations in deep variational quantum circuits, Advances in Neural Information Processing Systems35, 18612 (2022)
2022
-
[14]
Crognaletti, M
G. Crognaletti, M. Grossi, and A. Bassi, Estimates of loss function concentration in noisy parametrized quantum circuits, PRX Quantum7, 020336 (2026)
2026
-
[15]
S. H. Sack, R. A. Medina, A. A. Michailidis, R. Kueng, and M. Serbyn, Avoiding barren plateaus using classical shadows, PRX Quantum3, 020365 (2022)
2022
-
[16]
Y. Wang, B. Qi, C. Ferrie, and D. Dong, Trainabil- ity enhancement of parameterized quantum circuits via reduced-domain parameter initialization, Physical Re- view Applied22, 054005 (2024)
2024
-
[17]
Sannia, F
A. Sannia, F. Tacchino, I. Tavernelli, G. L. Giorgi, and R. Zambrini, Engineered dissipation to mitigate barren plateaus, npj Quantum Information10, 81 (2024)
2024
-
[18]
C.-Y. Park, M. Kang, and J. Huh, Hardware-efficient ansatz without barren plateaus in any depth, arXiv preprint arXiv:2403.04844 (2024)
2024 arXiv
-
[19]
Zapusek, I
E. Zapusek, I. Rojkov, and F. Reiter, Scaling quantum algorithms via dissipation: Avoiding barren plateaus, arXiv preprint arXiv:2507.02043 (2025)
2025 arXiv
-
[20]
Cichy, P
S. Cichy, P. K. Faehrmann, S. Khatri, and J. Eisert, Perturbative gadgets for gate-based quantum comput- ing: Nonrecursive constructions without subspace re- strictions, Physical Review A109, 052624 (2024)
2024
-
[21]
Gonz´ alez-Garc ´ ıa, J
G. Gonz´ alez-Garc ´ ıa, J. I. Cirac, and R. Trivedi, Pauli path simulations of noisy quantum circuits beyond aver- age case, Quantum9, 1730 (2025)
2025
-
[22]
Z. Chen, Y. Shao, Z. Liu, and Z. Wei, Taming bar- ren plateaus in arbitrary parameterized quantum cir- cuits without sacrificing expressibility, arXiv preprint arXiv:2511.13408 (2025)
2025
-
[23]
Cerezo, M
M. Cerezo, M. Larocca, D. Garc ´ ıa-Mart ´ ın, N. L. Diaz, P. Braccia, E. Fontana, M. S. Rudolph, P. Bermejo, A. Ijaz, S. Thanasilp,et al., Does provable absence of bar- ren plateaus imply classical simulability?, Nature Com- munications16, 7907 (2025)
2025
-
[24]
sympauli
the non-unitary channelE(σ) into a unitary chan- nelV(σ) acting on the enlarged Hilbert spaceH s ⊗ Ha, whereH s holds the system qubits andH a the ancillas. Initializing the ancillas in|0⟩and discarding them at the end reproduces the original channel, U(θ, σ) = Tra[V(σ)◦U(θ)(ρ...
2000
-
[25]
A. A. Mele, A. Angrisani, S. Ghosh, S. Khatri, J. Eisert, D. Stilck Fran¸ ca, and Y. Quek, Noise-induced shallow cir- cuits and the absence of barren plateaus, Nature Physics , 1 (2026)
2026
-
[26]
W. F. Stinespring, Positive functions on c*-algebras, Pro- ceedings of the American Mathematical Society6, 211 (1955)
1955
-
[27]
M. A. Nielsen and I. L. Chuang,Quantum Computation and Quantum Information(Cambridge University Press, 2000). 10
2000
-
[28]
A. W. Harrow and R. A. Low, Random quantum circuits are approximate 2-designs, Communications in Mathe- matical Physics291, 257 (2009)
2009
-
[29]
Kandala, A
A. Kandala, A. Mezzacapo, K. Temme, M. Takita, M. Brink, J. M. Chow, and J. M. Gambetta, Hardware- efficient variational quantum eigensolver for small molecules and quantum magnets, nature549, 242 (2017)
2017
-
[30]
J. Lee, W. J. Huggins, M. Head-Gordon, and K. B. Wha- ley, Generalized unitary coupled cluster wave functions for quantum computation, Journal of chemical theory and computation15, 311 (2018)
2018
-
[31]
Anand, P
A. Anand, P. Schleich, S. Alperin-Lea, P. W. Jensen, S. Sim, M. D ´ ıaz-Tinoco, J. S. Kottmann, M. Degroote, A. F. Izmaylov, and A. Aspuru-Guzik, A quantum com- puting view on unitary coupled cluster theory, Chemical Society Reviews51, 1659 (2022)
2022
-
[32]
R. Mao, P. Yuan, J. Allcock, and S. Zhang, Qaoa-maxcut has barren plateaus for almost all graphs, arXiv preprint arXiv:2512.24577 (2025)
2025
-
[33]
S. Kazi, M. Larocca, M. Farinati, P. J. Coles, M. Cerezo, and R. Zeier, Analyzing the quantum approximate op- timization algorithm: ans¨ atze, symmetries, and lie alge- bras, PRX Quantum6, 040345 (2025)
2025
-
[34]
K¨ okc¨ u, R
E. K¨ okc¨ u, R. Wiersema, A. F. Kemper, and B. N. Bakalov, Classification of dynamical lie algebras gener- ated by spin interactions on undirected graphs, arXiv preprint arXiv:2409.19797 (2024)
2024 arXiv
-
[35]
Bravyi and D
S. Bravyi and D. Gosset, Improved classical simulation of quantum circuits dominated by clifford gates, Physical review letters116, 250501 (2016)
2016
-
[36]
Aharonov, X
D. Aharonov, X. Gao, Z. Landau, Y. Liu, and U. Vazi- rani, A polynomial-time classical algorithm for noisy ran- dom circuit sampling, inProceedings of the 55th Annual ACM Symposium on Theory of Computing(2023) pp. 945–957
2023
-
[37]
Angrisani, A
A. Angrisani, A. Schmidhuber, M. S. Rudolph, M. Cerezo, Z. Holmes, and H.-Y. Huang, Classically esti- mating observables of noiseless quantum circuits, Physi- cal review letters135, 170602 (2025)
2025
-
[38]
Heyraud, Z
V. Heyraud, Z. Li, K. Donatella, A. Le Boit´ e, and C. Ciuti, Efficient estimation of trainability for vari- ational quantum circuits, PRX Quantum4, 040335 (2023)
2023
-
[39]
Y. Shao, Z. Chen, Z. Wei, and Z. Liu, Diagnosing quan- tum circuits: Noise robustness, trainability, and express- ibility, arXiv preprint arXiv:2509.11307 (2025)
2025
-
[40]
Javadi-Abhari, M
A. Javadi-Abhari, M. Treinish, K. Krsulich, C. J. Wood, J. Lishman, J. Gacon, S. Martiel, P. D. Nation, L. S. Bishop, A. W. Cross, B. R. Johnson, and J. M. Gambetta, Quantum computing with Qiskit (2024), arXiv:2405.08810 [quant-ph]
2024 arXiv
-
[41]
F. G. Brandao, M. Broughton, E. Farhi, S. Gutmann, and H. Neven, For fixed control parameters the quantum approximate optimization algorithm’s objective function value concentrates for typical instances, arXiv preprint arXiv:1812.04170 (2018)
2018 arXiv
-
[42]
Shaydulin, P
R. Shaydulin, P. C. Lotshaw, J. Larson, J. Ostrowski, and T. S. Humble, Parameter transfer for quantum approx- imate optimization of weighted maxcut, ACM Transac- tions on Quantum Computing4, 1 (2023)
2023
-
[43]
Mitarai, M
K. Mitarai, M. Negoro, M. Kitagawa, and K. Fujii, Quantum circuit learning, Physical Review A98, 032309 (2018)
2018
-
[44]
Wierichs, J
D. Wierichs, J. Izaac, C. Wang, and C. Y.-Y. Lin, General parameter-shift rules for quantum gradients, Quantum6, 677 (2022)
2022
-
[45]
ZHANG, T
Z. ZHANG, T. M. Ragonneau, and J. Schueller, zaikun- zhang/prima: Version 0.5 (2023)
2023
-
[46]
Y. Yan, M. Ma, Y. Zhou, and X. Ma, Variational locc- assisted quantum circuits for long-range entangled states, Physical Review Letters134, 170601 (2025)
2025
-
[47]
A. A. Mele, Introduction to haar measure tools in quan- tum information: A beginner’s tutorial, Quantum8, 1340 (2024)
2024
-
[48]
sympauli
R. Bhatia,Matrix Analysis, Vol. 169 (Springer, 1997). Appendix A: Methods In this section we present the proofs and theoretical analyses. We start with the proof of lemma 1. Proof.(Lemma 1) Recall thatU(θ, σ) is faithful toU(θ) if for the respective (H, ρ 0), the cost function...
1997
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.