REVIEW 3 major objections 6 minor 54 references
Regularized second-order optimization of tensor-network Born machines
T0 review · 3 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read This paper claims that log-smoothing regularization plus constrained Newton steps lets tensor-network Born machines converge faster and to lower negative log-likelihood than gradient descent or bare Newton's method.
desk verdict A genuinely new optimizer for TNBMs with consistent derivations, but the empirical evidence is thin and the key regularization-minima stability claim is untested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the constrained Newton step on the unit sphere of a single tensor, with the gradient and Hessian projected onto the tangent space so the step cannot leave the manifold of normalized states. The regularized loss replaces each overlap argument with $|(T,w_x)|^2+\epsilon$, which turns the poles into smooth shallow features; the optimizer then takes absolute values of all Hessian terms so the corrected Hessian is positive, a trust-region-like way to avoid converging to maxima. A decaying epsilon schedule starts from a smoothed landscape and gradually approaches the original loss, which is how the method claims to keep the benefits of smoothing without losing the original minima.
What would settle it
Take a fixed random initialization and run the regularized optimizer with several constant epsilon values, never decaying them, on the BAS dataset; then evaluate the unregularized NLL at each converged tensor. If the final unregularized loss changes systematically with epsilon, or if a line search from the regularized optimum along the direction to the original Newton optimum lowers the original loss, the smoothing has shifted the minima the optimizer is finding.
Extended reading notes
Core claim
The central claim is that the optimization bottleneck of TNBMs is not the model class but the geometry of the NLL cost: each training sample contributes a logarithmic pole at the hyperplane where its overlap vanishes, and these poles split the normalized-tensor manifold into about $2^{N_s}$ basins. The paper argues that a constrained Newton method on the hypersphere manifold, regularized by $l_\epsilon(x)=\log(|x|^2+\epsilon)$, can traverse those barriers, and that in the discrete BAS and MNIST benchmarks and the continuous Iris benchmark the regularized optimizer finds lower unregularized NLL than gradient descent or the unregularized constrained Newton method, with faster convergence and lower realization-to-realization variance. The claim includes a practical complexity story: the Newton step is solved with iterative Hessian-vector products, so the second-order update remains affordable.
Load-bearing premise
The method's reported gains depend on the assumption that replacing each log overlap with $\log(|x|^2+\epsilon)$ removes the problematic local minima without moving the remaining minima, so optimizing the smoothed loss is a faithful proxy for the original loss.
Editorial extensions
If this is right
- For the discrete benchmarks tested (BAS and MNIST), the regularized Newton optimizer reaches lower NLL after a full sweep than both steepest descent and unregularized Newton, with smaller variance across random initializations.
- On the continuous Iris benchmark, both regularization variants converge to low loss values, while bare Newton's method stalls at a high-loss local minimum; the bias variant converges sharply in the first sweep.
- The same algorithm can be applied to other tensor-network geometries such as PEPS and MERA by replacing the environment tensor, so the acceleration is not tied to one-dimensional MPS sweeps.
- The epsilon-regularized step costs only Hessian-vector products, so the per-iteration complexity stays $O(D N_s)$ for gradient and Hessian construction and roughly $O(D N_s N_{\mathrm{iter}})$ for the iterative linear solve, keeping second-order training practical for moderate bond dimensions.
Reading between the lines
- One testable extension the paper leaves open is an explicit check of the epsilon-schedule hypothesis: if the smoothing is doing the work, the final solution should be insensitive to the schedule within a broad range; if not, the reported gains may depend on the particular exponential decay used.
- The Lorentzian-convolution reading of the regularization suggests that other cheap smoothers of the log-overlap poles should behave similarly; a comparison of alternative kernels under the same Newton machinery would separate the smoothing mechanism from the specific epsilon-additive form.
- Outside TNBMs, the same combination of tangent-space Newton steps and log-singularity smoothing could be tried on any generative model whose likelihood has logarithmically diverging terms, such as tensor-network autoregressive or energy-based models, though the paper does not test those.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a regularized second-order optimization method for training tensor-network Born machines (TNBMs). It derives the projected gradient and Hessian of the single-tensor negative log-likelihood (NLL) loss on the manifold of normalized tensors, introduces a logarithmic smoothing regularization (and a bias-shift variant) to soften the singularities of the loss landscape, and modifies the Newton step with an absolute-value Hessian correction to avoid convergence to maxima. Experiments on BAS, MNIST, and Iris compare the method with gradient descent and unregularized constrained Newton, reporting lower unregularized NLL and faster convergence for the regularized method.
Significance. If the central claim holds, the paper offers a practical improvement to TNBM training: it shows how to combine on-manifold Newton steps with a simple regularization that smooths the log-pole landscape, and it provides explicit derivations of the regularized gradient and Hessian (Appendix A) and a useful complexity comparison (Table I). The experiments are reproducible in principle (five seeds, fixed schedules). The main weakness is that the bridge between the regularized loss actually minimized and the unregularized loss reported is asserted but not verified; the counterexample in this report shows the assertion is not generally true. The significance is therefore conditional on an additional empirical or theoretical check of minimum stability.
major comments (3)
- [III C, Eq. (16)] The regularized loss L_epsilon in Eq. (16) is what the optimizer actually minimizes, but the reported losses in Figs. 4 and 5 are the unregularized L0 from Eq. (4). The only bridge is the assertion that smoothing 'eliminates local minimum points' while 'the positions of the remaining minima remain largely stable.' This assertion is not verified quantitatively, and the convolution identity Eq. (18) applies only to the scalar logarithm, not to the coupled multivariate landscape. For a weighted sum of broadened logarithms, minima can shift substantially with epsilon: f_0(x) = -0.8 log|x| - 0.2 log|1-x| has its unique minimum at x=0.8, while the regularized version -0.8 log(x^2+eps) - 0.2 log((x-1)^2+eps) moves to x=0.2 for large eps. The TNBM NLL is exactly such a weighted sum over overlap zeros. The authors should provide evidence on BAS/MNIST/Iris with the chosen exponential schedule that the L_epsilon minimizers lie in basins of low L0, or temper the claim.
- [IV, Figs. 4-5] The experimental evaluation is limited to three small datasets (BAS, 7x7 MNIST, Iris) with five random seeds, a single learning rate for the gradient descent baseline (eta=0.05), and no sensitivity analysis for the regularization schedule (initial epsilon and decay rate). The abstract claims 'significantly enhances convergence rates and the quality of the optimized model' and calls the approach 'robust and scalable,' but the current evidence supports a more modest claim about the specific configurations tested. Please add hyperparameter robustness experiments or soften the conclusions.
- [Appendix A, Eq. (A7)] The absolute-value Hessian correction is introduced as a heuristic to 'eliminate the tendency of converging to local maximum points,' but no analysis is given of how this modification affects the Newton step or convergence. The analogy to trust-region methods is not exact: trust-region methods impose a provable bound on the step size, whereas here the step size is accepted without an explicit trust radius. This is load-bearing because the method's stability relies on this correction. Please provide at least an empirical study of the correction's effect or a theoretical justification.
minor comments (6)
- [III A] The notation for the constrained loss is inconsistent: Eq. (6) uses f(x) while the rest of the paper uses L(T); please unify the notation.
- [III A] Typo: 'Hesian' should be 'Hessian' in the paragraph following Eq. (14).
- [Figs. 2-3] Figures 2 and 3 lack axis labels and do not specify which tensor or direction is being sliced; please add these details.
- [Table I] In Table I, the symbols Ns, D, Niter, and kappa are not all defined in the caption; kappa appears in Niter ~ log(1/eps)*sqrt(kappa) but is never introduced.
- [IV] The description of the continuous Born machine relies entirely on reference [41]; please include the key details (feature embedding, isometric tensor update) so the Iris experiments can be reproduced.
- [V] The limitation paragraph in Section V is appreciated, but it does not mention the regularization-minimum shift issue; adding a caveat there or in Section III C would be helpful.
Circularity Check
No significant circularity: the regularized Newton optimizer is evaluated against independent baselines on the same objective, and the paper's central claim is not forced by its own definitions.
full rationale
The paper's target is the unregularized NLL of Eqs. (1) and (4), while the regularized loss L_eps of Eq. (16) is explicitly a surrogate with a smoothing parameter that is scheduled rather than fitted to the target. The claim that smoothing 'eliminates local minimum points' while 'the positions of the remaining minima remain largely stable' is a heuristic assertion about landscape geometry, not an identity: Eq. (18) only relates the scalar regularized log to a convolution and does not define the multivariate loss as equal to the original minimum. A reviewer concern that minima may shift with epsilon is an assumption-strength or correctness issue, not circularity, because the reported advantage is an empirical comparison against gradient descent and unregularized Newton on BAS, MNIST, and Iris, not an analytic consequence of defining the objective in terms of the method. The only self-referential element is reference [41] for the continuous Born machine architecture used in the Iris experiments; that reference supplies an experimental setup and is not invoked as a load-bearing uniqueness theorem or as the justification for the method's convergence claim. Section V further explicitly disclaims that the regularization mechanism is a general cure for local minima, which is inconsistent with the paper hiding a forced equivalence. No fitted parameter is renamed as a prediction, and no equation is shown to reduce to its own input by construction.
Assumptions & free parameters
free parameters (4)
- Regularization schedule initial epsilon =
0.025
- Gradient descent learning rate eta (continuous baseline) =
0.05
- MPS bond dimension chi =
5
- Feature dimensions for continuous Born machine =
25 to 3
assumptions (4)
- standard math Riemannian manifold optimization on the sphere (projected gradients and Hessians are correct)
- ad hoc to paper Regularization smooths logarithmic poles while preserving the positions of the remaining minima
- ad hoc to paper Absolute-value Hessian correction keeps the Newton step useful without converging to maxima
- domain assumption Continuous variables should map to same-sign amplitudes
Cite this review
Pith. "Pith review of Regularized second-order optimization of tensor-network Born machines." pith.science (2026). https://pith.science/paper/BDSSKH3J
@misc{pith2026250118691,
author = {Pith},
title = {Pith review of: Regularized second-order optimization of tensor-network Born machines},
year = {2026},
howpublished = {\url{https://pith.science/paper/BDSSKH3J}},
note = {Machine review of arXiv:2501.18691}
}
read the original abstract
Tensor-network Born machines (TNBMs) are quantum-inspired generative models for learning data distributions. Using tensor-network contraction and optimization techniques, the model learns an efficient representation of the target distribution, capable of capturing complex correlations with a compact parameterization. Despite their promise, the optimization of TNBMs presents several challenges. A key bottleneck of TNBMs is the logarithmic nature of the loss function commonly used for this problem. The single-tensor logarithmic optimization problem cannot be solved analytically, necessitating an iterative approach that slows down convergence and increases the risk of getting trapped in one of many non-optimal local minima. In this paper, we present an improved second-order optimization technique for TNBM training, which significantly enhances convergence rates and the quality of the optimized model. Our method employs a modified Newton's method on the manifold of normalized states, incorporating regularization of the loss landscape to mitigate local minima issues. We demonstrate the effectiveness of our approach by training a one-dimensional matrix product state (MPS) on both discrete and continuous datasets, showcasing its advantages in terms of stability and efficiency, and demonstrating its potential as a robust and scalable approach for optimizing quantum-inspired generative models.
Figures
Reference graph
Works this paper leans on
-
[41]
G. H. Golub, P. C. Hansen, and D. P. O’Leary, Tikhonov regularization and total least squares, SIAM journal on matrix analysis and applications 21, 185 (1999)
work page 1999
-
[1]
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, Generative adversarial nets, Advances in neural informa- tion processing systems 27 (2014)
work page 2014
-
[2]
J. Gui, Z. Sun, Y. Wen, D. Tao, and J. Ye, A review on generative adversarial networks: Algorithms, theory, and applications, IEEE transactions on knowledge and data engineering 35, 3313 (2021)
2021
-
[3]
and Boltzmann machines [4, 5], as well as transformer- based and diffusion-based models [6–9]. Among the diverse array of generative models, tensor- network Born machines (TNBMs) have emerged as a promising quantum-inspired framework for generative modeling [10–13]. TNBMs capitalize on the mathemat- ical framework of tensor networks to efficiently repre- ...
arXiv 2025
-
[4]
D. H. Ackley, G. E. Hinton, and T. J. Sejnowski, A learn- ing algorithm for boltzmann machines, Cognitive science 9, 147 (1985)
1985
-
[5]
D. P. Kingma, M. Welling, et al. , An introduction to variational autoencoders, Foundations and Trends ® in Machine Learning 12, 307 (2019)
work page 2019
-
[6]
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, At- tention is all you need, Advances in neural information processing systems 30 (2017)
work page 2017
-
[7]
R. Salakhutdinov and G. Hinton, Deep boltzmann ma- chines, in Artificial intelligence and statistics (PMLR,
Show all 54 references
-
[8]
J. Ho, A. Jain, and P. Abbeel, Denoising diffusion proba- bilistic models, Advances in neural information process- ing systems 33, 6840 (2020). 11
2020
-
[9]
T. Lin, Y. Wang, X. Liu, and X. Qiu, A survey of trans- formers, AI open 3, 111 (2022)
2022
-
[10]
Z.-Y. Han, J. Wang, H. Fan, L. Wang, and P. Zhang, Unsupervised generative modeling using matrix product states, Physical Review X 8, 031012 (2018)
2018
-
[11]
L. Yang, Z. Zhang, Y. Song, S. Hong, R. Xu, Y. Zhao, W. Zhang, B. Cui, and M.-H. Yang, Diffusion models: A comprehensive survey of methods and applications, ACM Computing Surveys 56, 1 (2023)
2023
-
[12]
Z.-Z. Sun, C. Peng, D. Liu, S.-J. Ran, and G. Su, Gener- ative tensor network classification model for supervised machine learning, Physical Review B101, 075135 (2020)
2020
-
[13]
Cheng, L
S. Cheng, L. Wang, T. Xiang, and P. Zhang, Tree tensor networks for generative modeling, Physical Review B 99, 155131 (2019)
2019
-
[14]
Verstraete, V
F. Verstraete, V. Murg, and J. I. Cirac, Matrix prod- uct states, projected entangled pair states, and varia- tional renormalization group methods for quantum spin systems, Advances in physics 57, 143 (2008)
2008
-
[15]
Cheng, L
S. Cheng, L. Wang, and P. Zhang, Supervised learning with projected entangled pair states, Physical Review B 103, 125117 (2021)
2021
-
[16]
Montangero, E
S. Montangero, E. Montangero, and Evenson, Introduc- tion to tensor network methods (Springer, 2018)
2018
-
[17]
Or´ us, A practical introduction to tensor networks: Matrix product states and projected entangled pair states, Annals of physics 349, 117 (2014)
R. Or´ us, A practical introduction to tensor networks: Matrix product states and projected entangled pair states, Annals of physics 349, 117 (2014)
2014
-
[18]
M. D. Garc ´ ıa and A. M. Romero, Survey on computa- tional applications of tensor-network simulations, IEEE Access (2024)
2024
-
[19]
M. C. Ba˜ nuls, Tensor network algorithms: A route map, Annual Review of Condensed Matter Physics 14, 173 (2023)
2023
-
[20]
M. C. Caro, H.-Y. Huang, M. Cerezo, K. Sharma, A. Sornborger, L. Cincio, and P. J. Coles, Generalization in quantum machine learning from few training data, Na- ture communications 13, 4919 (2022)
2022
-
[21]
Strashko and E
A. Strashko and E. M. Stoudenmire, Generalization and overfitting in matrix product state machine learning ar- chitectures, arXiv preprint arXiv: 2208.04372 (2022)
2022 arXiv
-
[22]
LeCun, C
Y. LeCun, C. Cortes, and C. Burges, Mnist hand- written digit database, ATT Labs [Online]. Available: http://yann.lecun.com/exdb/mnist 2 (2010)
2010
-
[23]
Schollw¨ ock, The density-matrix renormalization group in the age of matrix product states, Annals of physics 326, 96 (2011)
U. Schollw¨ ock, The density-matrix renormalization group in the age of matrix product states, Annals of physics 326, 96 (2011)
2011
-
[24]
Penrose, Applications of negative dimensional tensors, Combinatorial mathematics and its applications 1, 221 (1971)
R. Penrose, Applications of negative dimensional tensors, Combinatorial mathematics and its applications 1, 221 (1971)
1971
-
[25]
R. A. Fisher, Iris, UCI Machine Learning Repository (1988), DOI: https://doi.org/10.24432/C56C76
1988 doi
-
[26]
J. I. Cirac, D. Perez-Garcia, N. Schuch, and F. Ver- straete, Matrix product states and projected entangled pair states: Concepts, symmetries, theorems, Reviews of Modern Physics 93, 045003 (2021)
2021
-
[27]
Schuch, D
N. Schuch, D. P´ erez-Garc ´ ıa, and I. Cirac, Classi- fying quantum phases using matrix product states and projected entangled pair states, Physical Review B—Condensed Matter and Materials Physics 84, 165139 (2011)
2011
-
[28]
Catarina and B
G. Catarina and B. Murta, Density-matrix renormaliza- tion group: a pedagogical introduction, The European Physical Journal B 96, 111 (2023)
2023
-
[29]
G. K.-L. Chan and S. Sharma, The density matrix renor- malization group in quantum chemistry, Annual Review of Physical Chemistry 62, 465 (2011)
2011
-
[30]
Lubasch, J
M. Lubasch, J. I. Cirac, and M.-C. Banuls, Algorithms for finite projected entangled pair states, Physical Review B 90, 064425 (2014)
2014
-
[31]
Absil, R
P.-A. Absil, R. Mahony, and R. Sepulchre, Optimiza- tion algorithms on matrix manifolds (Princeton Univer- sity Press, 2008)
2008
-
[32]
Liao, J.-G
H.-J. Liao, J.-G. Liu, L. Wang, and T. Xiang, Differen- tiable programming tensor networks, Physical Review X 9, 031041 (2019)
2019
-
[33]
Lubasch, P
M. Lubasch, P. Moinier, and D. Jaksch, Multigrid renor- malization, Journal of Computational Physics 372, 587 (2018)
2018
-
[34]
J. J. Mor´ e and D. C. Sorensen, Computing a trust region step, SIAM Journal on scientific and statistical comput- ing 4, 553 (1983)
1983
-
[35]
Boumal, An introduction to optimization on smooth manifolds (Cambridge University Press, 2023)
N. Boumal, An introduction to optimization on smooth manifolds (Cambridge University Press, 2023)
2023
-
[36]
Yuan, A review of trust region algorithms for opti- mization, in Iciam, Vol
Y.-x. Yuan, A review of trust region algorithms for opti- mization, in Iciam, Vol. 99 (2000) pp. 271–282
2000
-
[37]
A. R. Conn, N. I. Gould, and P. L. Toint, Trust region methods (SIAM, 2000)
2000
-
[38]
bars” (columns), or “stripes
and the Tikhonov regularization procedure in regres- sion problems [39]. D. Second variant- regularization by shift A second version of regularization mechanism is rel- evant in cases when we can infer the preferable config- uration of amplitude signs a priori. In the case of ...
2023
-
[39]
Adachi, S
S. Adachi, S. Iwata, Y. Nakatsukasa, and A. Takeda, Solving the trust-region subproblem by a generalized eigenvalue problem, SIAM Journal on Optimization 27, 269 (2017)
2017
-
[40]
J. J. Mor´ e, The levenberg-marquardt algorithm: imple- mentation and theory, inNumerical analysis: proceedings of the biennial Conference held at Dundee, June 28–July 1, 1977 (Springer, 2006) pp. 105–116
1977
-
[42]
D. J. MacKay, Information theory, inference and learning algorithms (Cambridge university press, 2003)
2003
-
[43]
Meiburg, J
A. Meiburg, J. Chen, J. Miller, R. Tihon, G. Rabusseau, and A. Perdomo-Ortiz, Generative learning of continuous data by tensor networks, SciPost Physics 18, 096 (2025)
2025
-
[44]
Verstraete and J
F. Verstraete and J. I. Cirac, Renormalization algorithms for quantum-many body systems in two and higher di- mensions, arXiv preprint cond-mat/0407066 (2004)
2004 arXiv
-
[45]
Vidal, Entanglement renormalization, Physical review letters 99, 220405 (2007)
G. Vidal, Entanglement renormalization, Physical review letters 99, 220405 (2007)
2007
-
[46]
Evenbly and G
G. Evenbly and G. Vidal, Algorithms for entanglement renormalization, Physical Review B—Condensed Matter and Materials Physics 79, 144108 (2009)
2009
-
[47]
M. R. Hestenes, E. Stiefel, et al. , Methods of conjugate gradients for solving linear systems , Vol. 49 (NBS Wash- ington, DC, 1952)
1952
-
[48]
J. R. Shewchuk et al. , An introduction to the conjugate gradient method without the agonizing pain , Tech. Rep. (USA, 1994)
1994
-
[49]
C. C. Paige and M. A. Saunders, Solution of sparse in- definite systems of linear equations, SIAM journal on nu- merical analysis 12, 617 (1975)
1975
-
[50]
Saad, Iterative methods for sparse linear systems (SIAM, 2003)
Y. Saad, Iterative methods for sparse linear systems (SIAM, 2003)
2003
-
[51]
B. T. Polyak, Introduction to optimization (New York, Optimization Software,, 1987)
1987
-
[52]
Hackbusch, Iterative solution of large sparse systems of equations, Vol
W. Hackbusch, Iterative solution of large sparse systems of equations, Vol. 95 (Springer, 1994)
1994
-
[53]
Liesen and P
J. Liesen and P. Tich` y, Convergence analysis of krylov subspace methods, GAMM-Mitteilungen 27, 153 (2004)
2004
-
[54]
Haegeman, KrylovKit (2024)
J. Haegeman, KrylovKit (2024)
2024
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.