Pith. sign in

REVIEW 1 major objections 3 minor 35 references

Establishing Boundary KKT Convergence of Mirror Descent through Reparameterization

T0 review · 1 major / 3 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper proves that mirror descent converges to a KKT point even when the limit lies on the boundary, under joint conditions on the objective, the Legendre kernel, and the feasible geometry.

desk verdict A promising reparameterization framework with a real gap: the KKT conclusion of Theorem 4.1 is outsourced to an unstated companion-preprint proposition. read the letter →

arxiv 2608.07248 v1 pith:EAVJ2WZP submitted 2026-08-07 math.OC cs.LG

classification math.OCcs.LG MSC 90C3090C2649J52
keywords MirrordescentflowKurdyka–ŁojasiewiczinequalityCircuitdecompositionReparameterizationKKTstationarityMetricflatteningDefinableboundaryextension
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proves that mirror descent converges to a Karush–Kuhn–Tucker (KKT) point for smooth nonconvex problems with interval and affine constraints, even when the iterates accumulate on the boundary of the feasible set—the case where the Legendre gradient blows up and previous guarantees break down. The proof introduces a coordinatewise metric-flattening reparameterization $S_i(x_i)=\int \sqrt{h_i''(u)}\,du$ that turns the Hessian metric into the identity, so the boundary degeneracy disappears in the new variables. A conformal circuit decomposition bounds the Lagrangian gradient $\nabla f(x)+A^T\lambda$ uniformly up to the boundary, and the Kurdyka–Łojasiewicz inequality applied to the reparameterized objective gives finite length of $S(x_k)$. Continuity of the inverse reparameterization then recovers convergence of the original iterates to a KKT point. The framework covers Shannon entropy, Fermi–Dirac entropy, and power kernels on polyhedra.

What carries the argument

The workhorse is the metric-flattening reparameterization $S_i(x_i)=\int_{\bar x_i}^{x_i}\sqrt{h_i''(u)}\,du$, a coordinatewise change of variables with $D(S^{-1})(z)^T \nabla^2\phi(S^{-1}(z)) D(S^{-1})(z)=I$, so the Hessian metric becomes Euclidean and stays nondegenerate at the boundary. The second ingredient is a conformal circuit decomposition of the mirror-step displacement into elementary vectors of $\ker A$; it yields the boundary-uniform bound $\|\nabla f(x)+A^T\lambda\|_\infty\le\Gamma$, with $\Gamma$ independent of distance to the boundary. These two estimates feed the standard KL finite-length argument in $z$-coordinates, and the continuous definable extension of $S^{-1}$ carries convergence back to the original variables.

What would settle it

Locate and inspect Proposition 5.3 in the authors' earlier preprint [15]; if one can construct a convergent mirror-descent sequence with positive nonsummable stepsizes whose limit is not a first-order optimality point, the KKT conclusion of Theorem 4.1 collapses even though the finite-length part of the theorem could still hold.

Watch

Extended reading notes

Core claim

The central claim is that for separable Legendre kernels $\phi(x)=\sum_i h_i(x_i)$ with $h_i\in C^3$, $h_i''>0$, and the curvature bound $\kappa=\max_i \sup_t |h_i'''(t)|/(h_i''(t))^2<\infty$, mirror descent with stepsizes $0<\alpha_k\le\bar\alpha$, $\bar\alpha L<1$, and $\sum_k\alpha_k=\infty$ converges to a KKT point, provided the metric-flattening map admits a definable boundary extension. In detail, the reparameterized sequence has finite length, $S(x_k)\to z_*$ with $0\in\partial E(z_*)$, and $x_k\to S^{-1}(z_*)$, a point satisfying the KKT system $0\in\nabla f(x_*)+A^T\lambda_*+N_C(x_*)$. The paper also shows the same reparameterization converts the mirror flow into the Euclidean subgradient flow of $E$, giving a parallel trajectory theorem, and derives coordinatewise rates showing boundary coordinates can converge faster than interior ones.

Load-bearing premise

The load-bearing premise is an external quoted result, not proved in this manuscript, that any convergent mirror-descent sequence with positive nonsummable stepsizes must settle at a point satisfying the first-order optimality conditions.

Editorial extensions

If this is right

  • For Shannon-entropy mirror descent on a compact polytope, iterates converge to a KKT point with no interiority assumption on the limit.
  • The reparameterized sequence has finite length, so mirror descent converges as a sequence rather than only along subsequences.
  • Under a Kurdyka–Łojasiewicz exponent condition, rates transfer: finite termination for $\theta<1/2$, exponential for $\theta=1/2$, polynomial for $\theta>1/2$, with boundary coordinates converging faster than interior coordinates.
  • The same reparameterization gives finite length and KKT convergence for the continuous mirror flow.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the square-root reparameterization for Shannon entropy predicts that active (zero) coordinates contract at the squared rate of the reparameterized error, which could serve as an early stopping or active-set detection criterion.
  • Editorial inference: since the conformal-circuit bound only uses the null-space structure of the affine constraint, it may transfer to Bregman proximal point methods and Bregman ADMM, whose updates also lie in $\ker A$.
  • Editorial inference: the known non-KKT counterexample must violate at least one of the joint conditions (curvature bound, definable boundary extension, or relative smoothness); identifying which condition fails would delimit the true scope of the theorem.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 3 minor

Summary. The paper studies mirror descent for the nonconvex problem min{f(x): x∈C∩L}, where C is a product of open intervals and L is an affine space, using a separable Legendre kernel φ=Σh_i. It introduces a coordinatewise metric-flattening reparameterization S_i(x_i)=∫√h_i'' du and a reparameterized objective E(z)=f(S^{-1}(z)) on N=cl S(X°). Under conditions that κ=sup|(1/h_i'')'| is finite, that S admits a definable boundary extension, and that f is definable and relatively smooth, it proves that the reparameterized sequence S(x_k) has finite length and converges, that x_k converges to some x_*, and that 0∈∂E(z_*). The final KKT stationarity of x_* is asserted by invoking [15, Proposition 5.3] from the authors' previous preprint. The proof uses a conformal circuit decomposition to bound the Lagrangian gradient uniformly, one-step estimates transferring descent to the reparameterized variables, and a Kurdyka–Łojasiewicz finite-length argument with variable stepsizes. Three kernel families are treated: Shannon entropy, Fermi–Dirac entropy, and power kernels, with conditional coordinatewise convergence rates.

Significance. The paper addresses a real gap in nonconvex mirror descent theory: excluding boundary limits or regularizing the problem is standard, and a recent counterexample shows that non-KKT boundary accumulation is possible in general. The reparameterization framework and the circuit-based uniform bound on the Lagrangian gradient are novel and, together with the variable-stepsize KL argument, form an elegant proof of finite length in the reparameterized variables. The coverage of Shannon entropy, Fermi–Dirac entropy, and power kernels with rates is a concrete strength, as is the explicit verification of the definable boundary extension for these kernels. If the external KKT step is made self-contained, the main theorem would constitute a significant contribution.

major comments (1)
  1. [Section 4.2.3, proof of Theorem 4.1] The final step of the proof of Theorem 4.1 states 'Since {x_k} is convergent, [15, Proposition 5.3] implies that x_* is a KKT point of (1).' This is the only argument for the central KKT conclusion, but [15, Proposition 5.3] is neither stated nor proved in this manuscript. The preceding results only yield z_k→z_* with 0∈∂E(z_*), which is strictly weaker than KKT stationarity in the original variables because the Jacobian D(S^{-1})(z_*) can be degenerate at boundary coordinates (e.g., (S^{-1})'(0)=0 for Shannon entropy). The manuscript must state the full hypotheses of [15, Proposition 5.3] and either prove it in an appendix or verify in detail that all its hypotheses hold in the setting of Theorem 4.1, or provide a direct proof of KKT stationarity. As written, the main theorem is not self-contained and the central guarantee rests on an external result from the authors' own preprint.
minor comments (3)
  1. [Abstract and Section 1 (Contributions)] The phrase 'Continuity of S^{-1} then recovers convergence to the KKT point' is imprecise; continuity of S^{-1} recovers convergence of the original iterates, while KKT stationarity is supplied by the external proposition. Please rephrase to avoid the implication that KKT follows directly from continuity.
  2. [Appendix B] The global existence of the mirror flow relies on [15, Proposition 3.1] without a statement; a brief statement or proof would improve self-containedness, although this is not load-bearing for the discrete-time result.
  3. [Section 5.3] For p=1 the power kernel reduces to the Shannon kernel, but Proposition 3.3(iii) gives the closed-form formula only for p∈(1,2); the implicit definition through h''(x)=x^{-1} is workable, yet it is worth an explicit sentence noting the identification.

Circularity Check

1 steps flagged · score 5.0 of 10

The KKT conclusion of Theorem 4.1 is delegated to a same-author proposition [15, Prop. 5.3] that is neither stated nor proved; the paper's own argument only yields 0∈∂E(z*), which is weaker than KKT when D(S^{-1}) degenerates at the boundary.

  1. self citation load bearing [Section 4.2.3, Proof of Theorem 4.1 (final sentence); also Section 1.1]
    "Since {x_k} is convergent, [15, Proposition 5.3] implies that x_* is a KKT point of (1). ... Under the standing constraint qualification, it follows from [15, Proposition 5.3] that every convergent sequence generated by (3), with positive nonsummable stepsizes, has a KKT limit."

    The proof establishes finite length of S(x_k), z_k→z*, and 0∈∂E(z*), but this stationarity is not shown to imply KKT in the original variables. At a boundary coordinate, D(S^{-1})(z*)_{ii}=1/sqrt(h''_i(x*_i)) can vanish (e.g., Shannon entropy has (S^{-1})'(0)=0), so 0∈∂E(z*) imposes no sign condition on ∇f(x*)+A^Tλ. The only bridge from convergence-plus-stationarity to the theorem's central KKT conclusion is the quoted citation to [15, Prop. 5.3], a preprint by the same authors whose hypotheses are not stated or verified here and whose assertion is essentially the hard 'convergent ⇒ KKT' transfer being invoked. Thus the central guarantee is completed by an unproved, load-bearing self-citation rather than by the derivation in this manuscript.

full rationale

The metric-flattening reparameterization, conformal circuit bound (Lemma 4.8), one-step estimates (Prop. 4.11), and KL finite-length lemma (Lemma 4.12) are internally consistent, parameter-free, and not fitted to data; no fitted-input or definitional circularity is present. The convergence part of Theorem 4.1 is therefore a genuine independent contribution. However, the final KKT transfer is not derived: the proof ends by citing [15, Prop. 5.3], a same-author prior preprint that is not stated, proved, or machine-checked in this paper. Since 0∈∂E(z*) alone does not imply KKT when the reparameterization degenerates at the boundary, this self-citation is load-bearing for the paper's central claim. The same pattern appears in Theorem 4.3, where global existence of the mirror flow is imported from [15, Prop. 3.1]. The score is elevated above 2 because the central KKT guarantee is not self-contained, but it is not 8-10 because the reparameterized finite-length and convergence analysis is a nontrivial, independently derived result.

Assumptions & free parameters 0 free parameters · 8 assumptions · 0 invented entities

The new objects are the reparameterization S and the circuit seminorm, both defined from problem data. There are no free parameters fitted to data and no invented physical entities. The analysis assumes standard definability/KL machinery and one external proposition from the authors' prior work.

assumptions (8)
  • domain assumption Each h_i is a C3 Legendre kernel on I_i with h_i'' > 0
    Defines the kernel class throughout Sections 3-4 and Theorem 4.1.
  • domain assumption S admits a definable boundary extension (Definition 3.1)
    Gives compact N, a continuous definable inverse on the boundary, and the KL property for E; assumed in Theorem 4.1.
  • domain assumption Curvature bound κ = max_i sup_t |(1/h_i'')'(t)| < ∞
    Used in Lemma 4.10 and Proposition 4.11 to make the relative-error estimate independent of boundary distance.
  • domain assumption f is C1 on a neighborhood of X and L-smooth relative to φ on X°
    Assumed in Theorem 4.1; yields the sufficient-decrease Lemma 2.6.
  • domain assumption f|_X is definable in the common o-minimal structure
    Ensures the reparameterized objective E satisfies the KL property.
  • domain assumption Stepsizes satisfy 0 < α_k ≤ ᾱ < ∞, ᾱL < 1, and ∑ α_k = ∞
    Assumed in Theorem 4.1; needed for descent and KL summation.
  • standard math Conformal circuit decomposition (Lemma 4.5, after [25,30])
    Applied to decompose mirror-step displacements for the uniform bound on the Lagrangian gradient.
  • domain assumption Proposition 5.3 of [15]: convergent mirror-descent sequences with positive nonsummable stepsizes have KKT limits
    Invoked at the end of Theorem 4.1's proof to upgrade sequence convergence to KKT stationarity; not proved here.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Establishing Boundary KKT Convergence of Mirror Descent through Reparameterization." pith.science (2026). https://pith.science/paper/EAVJ2WZP

@misc{pith2026260807248,
  author       = {Pith},
  title        = {Pith review of: Establishing Boundary KKT Convergence of Mirror Descent through Reparameterization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EAVJ2WZP}},
  note         = {Machine review of arXiv:2608.07248}
}
abstract

We prove that mirror descent converges to a KKT point for the nonconvex problem without excluding boundary limits. The result holds under verifiable conditions that jointly couple the objective, the Legendre kernel, and the feasible geometry. The key ingredient to establish the convergence is a metric-flattening reparameterization \(S\) that admits a definable boundary extension. Applying the KL argument to the reparameterized objective yields convergence of \(S(x_k)\). Continuity of \(S^{-1}\) then recovers convergence to the KKT point of the original sequence. We further apply our general framework to some concrete examples: Shannon entropy, Fermi--Dirac entropy, and power kernels. Future work may consider more general constraint geometries and genuinely nonseparable kernels, and extend mirror descent to broader Bregman-type methods, e.g. Bregman proximal point algorithms and Bregman ADMM, and their inexact variants.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

35 extracted references · 27 canonical work pages

  1. [1]

    Hessian Riemannian gradient flows in convex programming

    Felipe Alvarez, J´ erˆ ome Bolte, and Olivier Brahic. Hessian Riemannian gradient flows in convex programming. SIAM Journal on Control and Optimization, 43(2):477–501, 2004

  2. [2]

    Singular Riemannian barrier methods and gradient-projection dynamical systems for constrained optimization.Optimization, 53(5–6):435–454, 2004

    Hedy Attouch, J´ erˆ ome Bolte, Patrick Redont, and Marc Teboulle. Singular Riemannian barrier methods and gradient-projection dynamical systems for constrained optimization.Optimization, 53(5–6):435–454, 2004

  3. [3]

    Hedy Attouch, J´ erˆ ome Bolte, and Benar Fux Svaiter. Convergence of descent methods for semi-algebraic and tame problems: Proximal algorithms, forward–backward splitting, and regularized Gauss–Seidel methods.Mathematical Programming, 137:91–129, 2013. 21

  4. [4]

    Bauschke, J´ erˆ ome Bolte, and Marc Teboulle

    Heinz H. Bauschke, J´ erˆ ome Bolte, and Marc Teboulle. A descent lemma beyond Lipschitz gradient continuity: First-order methods revisited and applications.Mathematics of Operations Research, 42(2):330–348, 2017

  5. [5]

    J´ erˆ ome Bolte, Aris Daniilidis, and Adrian S. Lewis. The Lojasiewicz inequality for nonsmooth subanalytic functions with applications to subgradient dynamical systems.SIAM Journal on Optimization, 17(4):1205–1223, 2007

  6. [6]

    Lewis, and Masahiro Shiota

    J´ erˆ ome Bolte, Aris Daniilidis, Adrian S. Lewis, and Masahiro Shiota. Clarke subgradients of stratifiable functions. SIAM Journal on Optimization, 18(2):556–572, 2007

  7. [7]

    Curiosities and counterexamples in smooth convex optimization.Mathematical Programming, 195(1–2):553–603, 2022

    J´ erˆ ome Bolte and Edouard Pauwels. Curiosities and counterexamples in smooth convex optimization.Mathematical Programming, 195(1–2):553–603, 2022

  8. [8]

    First order methods beyond convexity and Lipschitz gradient continuity with applications to quadratic inverse problems.SIAM Journal on Optimization, 28(3):2131–2151, 2018

    J´ erˆ ome Bolte, Shoham Sabach, Marc Teboulle, and Yakov Vaisbourd. First order methods beyond convexity and Lipschitz gradient continuity with applications to quadratic inverse problems.SIAM Journal on Optimization, 28(3):2131–2151, 2018

Show all 35 references
  1. [9]

    Bomze, Panayotis Mertikopoulos, Werner Schachinger, and Mathias Staudigl

    Immanuel M. Bomze, Panayotis Mertikopoulos, Werner Schachinger, and Mathias Staudigl. Hessian barrier algorithms for linearly constrained optimization problems.SIAM Journal on Optimization, 29(3):2100–2127, 2019

  2. [10]

    Spurious stationarity and hardness results for Bregman proximal- type algorithms.arXiv preprint arXiv:2404.08073, 2026

    He Chen, Jiajin Li, and Anthony Man-Cho So. Spurious stationarity and hardness results for Bregman proximal- type algorithms.arXiv preprint arXiv:2404.08073, 2026

  3. [11]

    On the iterate convergence of Bregman projected gradient method.arXiv preprint arXiv:2608.05035, 2026

    He Chen and Anthony Man-Cho So. On the iterate convergence of Bregman projected gradient method.arXiv preprint arXiv:2608.05035, 2026

  4. [12]

    Sinkhorn distances: Lightspeed computation of optimal transport

    Marco Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport. InAdvances in Neural Information Processing Systems 26, 2013

  5. [13]

    Dang and Guanghui Lan

    Cong D. Dang and Guanghui Lan. On the convergence properties of non-Euclidean extragradient methods for variational inequalities with generalized monotone operators.Computational Optimization and Applications, 60(2):277–310, 2015

  6. [14]

    Nonconvex stochastic Bregman proximal gradient method with application to deep learning.Journal of Machine Learning Research, 26(39):1–44, 2025

    Kuangyu Ding, Jingyang Li, and Kim-Chuan Toh. Nonconvex stochastic Bregman proximal gradient method with application to deep learning.Journal of Machine Learning Research, 26(39):1–44, 2025

  7. [15]

    On exploration of an interior mirror descent flow for stochastic nonconvex constrained problem.arXiv preprint arXiv:2507.15264v3, 2025

    Kuangyu Ding and Kim-Chuan Toh. On exploration of an interior mirror descent flow for stochastic nonconvex constrained problem.arXiv preprint arXiv:2507.15264v3, 2025

  8. [16]

    Non-KKT accumulation in entropic mirror descent, 2026

    Kuangyu Ding and Kim-Chuan Toh. Non-KKT accumulation in entropic mirror descent, 2026. arXiv:2608.01658

  9. [17]

    Doan, Subhonmesh Bose, D

    Thinh T. Doan, Subhonmesh Bose, D. Hoa Nguyen, and Carolyn L. Beck. Convergence of the iterates in mirror descent methods.IEEE Control Systems Letters, 3(1):114–119, 2019

  10. [18]

    A bregman ADMM for Bethe variational problem.arXiv preprint arXiv:2502.04613, 2025

    Yuehaw Khoo, Tianyun Tang, and Kim-Chuan Toh. A bregman ADMM for Bethe variational problem.arXiv preprint arXiv:2502.04613, 2025

  11. [19]

    On gradients of functions definable in O-minimal structures.Annales de l’Institut Fourier, 48(3):769–783, 1998

    Krzysztof Kurdyka. On gradients of functions definable in O-minimal structures.Annales de l’Institut Fourier, 48(3):769–783, 1998

  12. [20]

    A convergent single-loop algorithm for relaxation of Gromov–Wasserstein in graph data

    Jiajin Li, Jianheng Tang, Lemin Kong, Huikang Liu, Jia Li, Anthony Man-Cho So, and Jose Blanchet. A convergent single-loop algorithm for relaxation of Gromov–Wasserstein in graph data. InInternational Conference on Learning Representations, 2023

  13. [21]

    Convergence of the exponentiated gradient method with Armijo line search

    Yen-Huan Li and Volkan Cevher. Convergence of the exponentiated gradient method with Armijo line search. Journal of Optimization Theory and Applications, 181(2):588–607, 2019

  14. [22]

    Lee, and Sanjeev Arora

    Zhiyuan Li, Tianhao Wang, Jason D. Lee, and Sanjeev Arora. Implicit bias of gradient descent on reparametrized models: On equivalence to mirror descent. InAdvances in Neural Information Processing Systems 35, pages 34626–34640, 2022

  15. [23]

    Gromov–Wasserstein distances and the metric approach to object matching.Foundations of Computational Mathematics, 11(4):417–487, 2011

    Facundo M´ emoli. Gromov–Wasserstein distances and the metric approach to object matching.Foundations of Computational Mathematics, 11(4):417–487, 2011

  16. [24]

    Global convergence of model function based Bregman proximal minimization algorithms.Journal of Global Optimization, 83(4):753–781, 2022

    Mahesh Chandra Mukkamala, Jalal Fadili, and Peter Ochs. Global convergence of model function based Bregman proximal minimization algorithms.Journal of Global Optimization, 83(4):753–781, 2022

  17. [25]

    Elementary vectors and conformal sums in polyhedral geometry and their relevance for metabolic pathway analysis.Frontiers in Genetics, 7:90, 2016

    Stefan M¨ uller and Georg Regensburger. Elementary vectors and conformal sums in polyhedral geometry and their relevance for metabolic pathway analysis.Frontiers in Genetics, 7:90, 2016

  18. [26]

    Computational optimal transport.Foundations and Trends in Machine Learning, 11(5–6):355–607, 2019

    Gabriel Peyr´ e and Marco Cuturi. Computational optimal transport.Foundations and Trends in Machine Learning, 11(5–6):355–607, 2019. 22

  19. [27]

    Gromov–Wasserstein averaging of kernel and distance matrices

    Gabriel Peyr´ e, Marco Cuturi, and Justin Solomon. Gromov–Wasserstein averaging of kernel and distance matrices. InProceedings of the 33rd International Conference on Machine Learning, volume 48 ofProceedings of Machine Learning Research, pages 2664–2672, 2016

  20. [28]

    Shuffling the stochastic mirror descent via dual lipschitz continuity and kernel conditioning.arXiv preprint arXiv:2603.16042, 2026

    Junwen Qiu, Leilei Mei, and Junyu Zhang. Shuffling the stochastic mirror descent via dual lipschitz continuity and kernel conditioning.arXiv preprint arXiv:2603.16042, 2026

  21. [29]

    Entropic Gromov–Wasserstein distances: Stability and algorithms

    Gabriel Rioux, Ziv Goldfeld, and Kengo Kato. Entropic Gromov–Wasserstein distances: Stability and algorithms. Journal of Machine Learning Research, 25(363):1–52, 2024

  22. [30]

    Tyrrell Rockafellar

    R. Tyrrell Rockafellar. The elementary vectors of a subspace of RN . In R. C. Bose and T. A. Dowling, editors, Combinatorial Mathematics and Its Applications, pages 104–127. University of North Carolina Press, Chapel Hill, NC, 1969

  23. [31]

    Tyrrell Rockafellar.Convex Analysis, volume 28 ofPrinceton Mathematical Series

    R. Tyrrell Rockafellar.Convex Analysis, volume 28 ofPrinceton Mathematical Series. Princeton University Press, Princeton, NJ, 1970

  24. [32]

    Tyrrell Rockafellar and Roger J.-B

    R. Tyrrell Rockafellar and Roger J.-B. Wets.Variational Analysis. Springer, Berlin, 1998

  25. [33]

    Linear-time Gromov–Wasserstein distances using low-rank couplings and costs

    Meyer Scetbon, Gabriel Peyr´ e, and Marco Cuturi. Linear-time Gromov–Wasserstein distances using low-rank couplings and costs. InProceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 19347–19365, 2022

  26. [34]

    On the convergence rate of stochastic mirror descent for nonsmooth nonconvex optimization

    Siqi Zhang and Niao He. On the convergence rate of stochastic mirror descent for nonsmooth nonconvex optimization. arXiv:1806.04781, 2018

  27. [35]

    Boyd, and Peter W

    Zhengyuan Zhou, Panayotis Mertikopoulos, Nicholas Bambos, Stephen P. Boyd, and Peter W. Glynn. On the convergence of mirror descent beyond stochastic convex programming.SIAM Journal on Optimization, 30(1):687–716, 2020. 23

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.