Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Heavy Ball and Nesterov Accelerations with Hessian-driven Damping for Nonconvex Optimization

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper proves exponential convergence for a Hessian-damped ODE and linear convergence for its two discretized momentum algorithms on strongly quasiconvex nonconvex objectives.

desk verdict Mostly sound conditional theory for Hessian-damped momentum on strongly quasiconvex functions, but the numerical validation lies outside the proven parameter regimes and the proven Nesterov regime is too narrow to support the acceleration claims. read the letter →

arxiv 2506.15632 v1 pith:JQEDWDHF submitted 2025-06-18 math.OC

classification math.OC MSC 90C2665K1034A34
keywords stronglyquasiconvexHessian-drivendampingheavyballmethodNesterovaccelerationlinearconvergenceexponentialnonconvexoptimizationsecond-orderdynamicalsystem
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper studies minimization of differentiable strongly quasiconvex functions, a family that includes strongly convex functions along with genuinely nonconvex examples such as sub-unit norms and certain fractional ratios. It proposes a second-order dynamical system with both constant viscous damping and Hessian-driven damping, and proves that under a kappa-quasar-convexity inequality every trajectory converges exponentially to the unique minimizer. Discretizing that system produces a Heavy Ball method and a Nesterov-type method, each carrying a finite-difference Hessian-correction term; the paper proves linear convergence of both iterates and function values for each method. The point of the package is to bring accelerated momentum guarantees to a nonconvex class while using curvature information to suppress the oscillations that plain momentum produces.

What carries the argument

The load-bearing object is a Lyapunov energy function. For the continuous system the energy is $E(t)=h(x(t))-h_*+\frac{1}{2}\|\lambda(x(t)-\bar{x})+\dot{x}(t)+\beta\nabla h(x(t))\|^2$ with $\lambda=2\alpha/(\kappa+2)$, and its derivative is shown to satisfy $\dot{E}(t)+\frac{\lambda\kappa}{2}E(t)\le 0$. For the discrete methods the energy is $E_k=h(x_k)-h_*+\frac{\alpha^2}{\theta+\beta}\|x_k-x_{k-1}\|^2+\frac{\theta^2}{\theta+\beta}\|\nabla h(x_{k-1})\|^2$ for the Heavy Ball variant and a position-momentum energy for the Nesterov variant. Each proof combines the descent lemma for $L$-smooth functions with the differential characterization of strong quasiconvexity, gradient dominance, and Assumption (23). The Hessian-correction term enters only through gradient differences, so neither algorithm requires an actual Hessian evaluation.

What would settle it

Simulate the damped ODE (22) on the Example 9 function $h(x)=h_1(x)+x^2$, which is strongly quasiconvex but not quasar-convex: if the trajectory still decays at the claimed rate with the claimed constants, Assumption (23) is not necessary; if the decay is slower or fails, the theorem's scope is exactly as stated. For the discrete side, run the Heavy Ball method (39) with the Experiment 20 parameters $\alpha=0.8$, $\theta=0.05$, and $\beta=1/24$, and check whether the proof's contraction factor $1-\rho/\sigma$ is actually below one, since those parameters lie outside the box (40).

Watch

Extended reading notes

Core claim

The central discovery is that Hessian-driven damping, already known to improve heavy-ball dynamics for convex problems, supplies the same Lyapunov contraction for strongly quasiconvex objectives. Theorem 5 states that if $h$ is twice differentiable, strongly quasiconvex with modulus $\gamma>0$, and satisfies Assumption (23), then for $\alpha\in(0,\sqrt{\gamma(\kappa+2)^2/(8\kappa)}]$ and $\beta\in(0,(\kappa+2)/(\kappa\alpha))$ the trajectory of (22) satisfies $h(x(t))-h_*\le C e^{-\alpha\kappa t/(\kappa+2)}$. The discrete counterpart replaces the Hessian term $\nabla^2 h(x)\dot{x}$ by the gradient difference $\theta(\nabla h(x_k)-\nabla h(x_{k-1}))$; Theorem 13 gives linear convergence for the Heavy Ball form (39) under parameter restrictions (40), and Theorem 17 gives linear convergence for the Nesterov-type form (55) under condition (59). In both discrete theorems a Lyapunov energy contracts by a factor in $(0,1)$ at every step, so iterates and function values converge linearly to the unique minimizer.

Load-bearing premise

The continuous-time exponential rate rests on Assumption (23), a $\kappa$-quasar-convexity inequality that strong quasiconvexity alone does not imply; the paper itself gives an unbounded strongly quasiconvex function that violates it.

Editorial extensions

If this is right

  • For every twice differentiable strongly quasiconvex function satisfying Assumption (23), the damped trajectory (22) reaches its unique minimizer with function values obeying $h(x(t))-h_*\le C e^{-\alpha\kappa t/(\kappa+2)}$ from any initial condition.
  • The Heavy Ball method with Hessian correction (39) converges linearly for $L$-smooth strongly quasiconvex functions whenever the parameters satisfy condition (40).
  • The Nesterov-type method with adaptive momentum (55) converges linearly under condition (59), with contraction factor $\mu_1(1+\varepsilon)$.
  • Both discrete methods inherit oscillation suppression from the continuous Hessian damping because the term $\theta(\nabla h(x_k)-\nabla h(x_{k-1}))$ acts as a finite-difference surrogate for $\nabla^2 h(x_k)(x_k-x_{k-1})$.
  • Exponential and linear convergence are established without convexity; the classical strongly convex behavior is included as a special case of the strongly quasiconvex class.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the linear rates survive outside the parameter box (40), the same gradient-difference correction could be added to stochastic or proximal heavy-ball variants, giving curvature awareness without Hessian computations.
  • The sufficient conditions (35)-(36) turn strong quasiconvexity into quasar-convexity under quadratic-growth control; a natural next step is to relax Assumption (23) to a local or tail inequality and ask how much of the exponential rate is lost.
  • The continuous-time analysis suggests that Hessian-driven damping should also suppress oscillations for other nonconvex classes satisfying gradient dominance, such as Polyak-Łojasiewicz-type functions, a transfer the paper does not make.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper studies unconstrained minimization of strongly quasiconvex functions through a second-order ODE with Hessian-driven damping, and derives two discrete-time algorithms from it: a Heavy Ball method with Hessian correction (39) and a Nesterov-type method with adaptive momentum (55). For the ODE, Theorem 5 proves exponential convergence of function values and iterates under an additional quasar-convexity-type assumption (23). For the discrete methods, Theorem 13 proves linear convergence of the Heavy Ball iteration under parameter restriction (40), and Theorem 17 proves linear convergence of the Nesterov-type iteration under restriction (59). Section 3.2 analyzes sufficient conditions for assumption (23), including Propositions 10 and 11 and counterexamples. Section 5 gives numerical experiments on two nonconvex strongly quasiconvex functions. The paper concludes that Hessian-driven damping accelerates convergence and reduces oscillations relative to classical momentum methods.

Significance. If the results hold in the stated regimes, the paper would extend the known exponential/linear convergence guarantees for strongly convex functions and for the heavy ball ODE in [29] to a class of nonconvex strongly quasiconvex functions with Hessian damping, and it would provide two explicit discretizations with linear convergence. The Lyapunov arguments in Theorems 5, 13, and 17 are carefully executed, the algebra appears coherent under the stated parameter conditions, and the limiting cases θ=0 and β=0 connect cleanly to existing results in [29]. The counterexamples in Section 3.2 are valuable and show the authors are aware that assumption (23) is not automatic. However, the numerical experiments do not satisfy the hypotheses of the theorems, and the parameter regime allowed by Theorem 17 is narrow enough to undermine the practical claim of acceleration. The paper is a meaningful contribution to the theory of Hessian-damped momentum methods for generalized convex functions, but its advertised experimental support and its abstract-level claims need to be reconciled with the proven conditional statements.

major comments (4)
  1. [Section 5, Examples 20–21] The numerical experiments do not fall within the hypotheses of the theorems they are claimed to illustrate. In Example 20, the Heavy Ball run uses α=0.8, but Theorem 13 condition (40) requires α<√2/2≈0.707; moreover, with α=0.8 the quantity (1−2α²)/L is negative, so no positive θ+β can satisfy θ+β≤(1−2α²)/L. For the Nesterov-type run, Theorem 17 requires β<γ/(ηL²) for some η>1, which with γ=1/2 and L=6 gives β<1/72, while the experiment uses β=1/24≈0.0417. Example 21 similarly chooses α=0.8 and 0.9 with β=0.0025 and tuned θ without checking (40) or (59). Therefore the oscillation-reduction and linear-rate behavior shown in Figures 1 and 3 is not an instantiation of the proven theorems, and the abstract's statement that numerical experiments support the obtained results is not justified. The experiments should be rerun inside the proven parameter regimes or explicitly labeled as heuristic extensions outside the scope of the theorems.
  2. [Section 4.2, Theorem 17 and condition (59)] The proven parameter regime for the Nesterov-type method is so restrictive that the method is effectively gradient descent rather than an accelerated method. For the allowed range β<γ/(ηL²), the quantity µ1 is close to 1 and µ2 is at most 1−1/η, so condition (59) forces α+θL to be very small. For example, with γ=1/2, L=6, η=2 and β=1/150, the admissible ǫ is at most about 1.3×10⁻⁴ and (59) forces α+θL ≲ 7×10⁻⁵. In this regime the momentum coefficient α is negligible and the method is essentially gradient descent. The authors should state this limitation explicitly or find a less restrictive condition if the advertised acceleration is to be meaningful.
  3. [Section 3, Theorem 5 and Examples 9, 12] Theorem 5 is stated for strongly quasiconvex functions but its proof relies crucially on assumption (23), which is not implied by strong quasiconvexity. Example 9 (and also Example 12) explicitly constructs a strongly quasiconvex function that is not quasar-convex and hence need not satisfy (23). Since the theorem is conditional on (23), the abstract and introduction should state more carefully that the exponential-convergence result applies to strongly quasiconvex functions satisfying assumption (23), with Section 3.2 providing sufficient conditions. As written, the claim that the paper studies the dynamical system 'tailored for a class of nonconvex functions called strongly quasiconvex' overstates the domain of validity.
  4. [Section 4.1, proof of Theorem 13 around Eq. (50)] The proof divides by α²−3θ²L² when deriving inequality (50). This quantity is positive under condition (40) because θ<α/(L√3), but the division is done without explicitly noting the positivity. In the boundary case α=0 the condition θ∈[0, α/(L√3)) is empty, so the theorem statement should restrict α>0 or treat α=0 separately. This is a local clarity issue, but it should be fixed because an unnamed division by a potentially zero quantity is a correctness hazard in an otherwise sound proof.
minor comments (5)
  1. [Section 5, Example 20 and Figure 1] The text defines h(x)=x²+2sin²x, but the caption of Figure 1 states h(x)=x²+3sin²(x); these should be made consistent.
  2. [Section 5, Example 21] The sentence 'it can be verified that h it is strongly quasiconvex with modulus' is grammatically incomplete and omits the value of the modulus; 'parameters turning' should read 'parameter tuning'.
  3. [Corollary 7] The display in the proof of Corollary 7 is malformed: the fragment '˜C κ ∈ (2, +∞)' and the duplicated exponential-integral inequalities make the statement very hard to parse. The corollary should be restated with explicit constants or with a cleaner piecewise formulation.
  4. [Section 6] The concluding paragraph says 'future research lines are discussed in Section 5', but the conclusions and future directions are in Section 6.
  5. [Remark 8] The remark says Theorem 5 allows 'a free damping term', but Theorem 5 restricts α to the interval (0, √(γ(κ+2)²/(8κ))], so the damping parameter is not literally free; the wording should be softened.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: the convergence rates follow from explicit hypotheses via independent Lyapunov arguments; self-citations are contextual and not load-bearing.

full rationale

The paper's central claims are conditional convergence theorems. Theorem 5 assumes the kappa-quasar-convexity inequality (23) and proves exponential decay by a Lyapunov argument; the assumption is not the conclusion, and Section 3.2 independently derives sufficient conditions (Proposition 10 and Corollary 11) for strongly quasiconvex functions to satisfy it, using established characterizations (Lemmas 3 and 4). No parameter is fitted and then renamed as a prediction: the rates are explicit functions of the stated constants (alpha, beta, gamma, L, kappa). Theorems 13 and 17 are proved directly for the discrete iterations via energy estimates, using standard smoothness/strong-quasiconvexity/PL facts (Lemma 4 is cited to an external source, [27], and the quadratic-growth inequality (13) to [34]); these proofs do not import the continuous-time result as a premise. The self-citations to [28], [29], and [30] appear in contextual comparisons and numerical examples (e.g., Example 20's function is asserted strongly quasiconvex by [29, Example 30]) and are not load-bearing for the main convergence proofs. Numerical parameter choices in Examples 20-21 fall outside the theorem hypotheses, but that is a consistency/correctness issue, not circularity. No step in the derivation reduces, by definition or by self-citation, to the result being proven.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The theoretical results have no fitted constants; the parameters alpha, theta, and beta are algorithm hyperparameters, listed here because they are hand-chosen in the experiments and, in those experiments, chosen outside the proven ranges. The main extra assumption beyond strong quasiconvexity is the quasar-convexity inequality (23), plus unstated ODE well-posedness.

free parameters (3)
  • alpha (momentum coefficient) = HB: 0.8; Nesterov: 0.6; Example 21: 0.8/0.9
    Chosen by hand in the numerical experiments; the HB value 0.8 violates Theorem 13's requirement alpha < 1/sqrt(2), and the Nesterov choices are not checked against condition (59).
  • theta (Hessian correction coefficient) = 0.05; Example 21: 0.004/0.009
    Hand-tuned without derivation from L or gamma; in Example 20, theta=0.05 with alpha=0.8 cannot satisfy the conditions of Theorem 13.
  • beta (step size) = 1/(4L)=1/24; Example 21: 0.0025
    Set by hand; beta=1/24 exceeds gamma/L^2=1/72 for Example 20, so the Nesterov theorem's condition beta < gamma/(eta L^2) cannot hold for any eta>1.
assumptions (5)
  • domain assumption Strong quasiconvexity definition (12) and Lemma 1: unique minimizer with quadratic growth (13).
    Basis for the function class; cited from prior literature, not proved in the paper.
  • ad hoc to paper Assumption (23): <nabla h(x), x - xbar> >= kappa (h(x) - h(xbar)).
    Imposed for Theorem 5; not implied by strong quasiconvexity alone, as Example 9 shows, though L-smoothness gives kappa = gamma/L.
  • standard math L-smoothness (15) and the descent lemma (16).
    Used throughout the discrete proofs and in deriving the PL inequality.
  • domain assumption PL inequality of Lemma 4 with modulus mu = gamma^2/(2L).
    Quoted from [27]; key to bounding the discrete Lyapunov function and depends on L-smoothness plus strong quasiconvexity.
  • ad hoc to paper Existence and uniqueness of solutions to the ODE (22) for twice differentiable h.
    The paper does not prove well-posedness under its assumptions; it treats trajectories x(t) as given, which is a load-bearing unstated premise.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Heavy Ball and Nesterov Accelerations with Hessian-driven Damping for Nonconvex Optimization." pith.science (2026). https://pith.science/paper/JQEDWDHF

@misc{pith2026250615632,
  author       = {Pith},
  title        = {Pith review of: Heavy Ball and Nesterov Accelerations with Hessian-driven Damping for Nonconvex Optimization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JQEDWDHF}},
  note         = {Machine review of arXiv:2506.15632}
}
read the original abstract

In this work, we investigate a second-order dynamical system with Hessian-driven damping tailored for a class of nonconvex functions called strongly quasiconvex. Buil\-ding upon this continuous-time model, we derive two discrete-time gra\-dient-based algorithms through time discretizations. The first is a Heavy Ball method with Hessian correction, incorporating cur\-va\-tu\-re-dependent terms that arise from discretizing the Hessian damping component. The second is a Nesterov-type accelerated method with adaptive momentum, fea\-tu\-ring correction terms that account for local curvature. Both algorithms aim to enhance stability and convergence performance, particularly by mi\-ti\-ga\-ting oscillations commonly observed in cla\-ssi\-cal momentum me\-thods. Furthermore, in both cases we establish li\-near convergence to the optimal solution for the iterates and functions values. Our approach highlights the rich interplay between continuous-time dynamics and discrete optimization algorithms in the se\-tting of strongly quasiconvex objectives. Numerical experiments are presented to support obtained results.

Figures

Figures reproduced from arXiv: 2506.15632 by the authors.

Figure 1
Figure 1. Comparison in objective values (left) and trajectories (r [PITH_FULL_IMAGE:figures/full_fig_p023_1.png] view at source ↗
Figure 2
Figure 2. An illustration of the strongly quasiconvex function [PITH_FULL_IMAGE:figures/full_fig_p024_2.png] view at source ↗
Figure 3
Figure 3. Comparison in objective values (left) and trajectories (r [PITH_FULL_IMAGE:figures/full_fig_p024_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Star Quasiconvexity: a Unified Approach for Linear Convergence of First-Order Methods Beyond Convexity

    math.OC 2025-10 conditional novelty 4.0 of 10

    A function class whose sublevel sets are star-shaped at a minimizer unifies several generalized convexity conditions and yields linear convergence of first-order methods, including a new proximal point result on star-...

Reference graph

Works this paper leans on

45 extracted references · 38 canonical work pages · cited by 1 Pith paper

  1. [29]

    Lara, R.T

    F. Lara, R.T. Marcavillaca, P.T. Vuong, Characterizations, dynam- ical systems and gradient methods for strongly quasiconvex func tions, J. Optim. Theory Appl. , 206, DOI 10.1007/s10957-025-02728-y (2025)

  2. [1]

    Antipin, Minimization of convex functions on convex sets by means of differential equations, Diff

    A.S. Antipin, Minimization of convex functions on convex sets by means of differential equations, Diff. Eq. , 30, 1365–1375, (1994)

  3. [2]

    Alvarez, On the minimizing property of a second order dissipative dynamical system in Hilbert spaces, SIAM J

    F. Alvarez, On the minimizing property of a second order dissipative dynamical system in Hilbert spaces, SIAM J. Control Optim. , 39, 1102– 1119, (2000)

  4. [3]

    Alvarez, H

    F. Alvarez, H. Attouch, J. Bolte, P. Redont , A second-order gradient-like dissipative dynamical system with Hessian-driven damp ing: application to optimization and mechanics, J. Math. Pures et Appl. , 81(3), 747–779, (2002)

  5. [4]

    Attouch, Z

    H. Attouch, Z. Chbani, J. Fadili, H. Riahi , First-order optimization algorithms via inertial systems with Hessian driven damping, Math. Pro- gram., 193, 113–155, (2022). 25

  6. [5]

    Attouch, X

    H. Attouch, X. Goudon, P. Redont , The heavy ball with friction I The continuous dynamical system, Commun. Contemp. Math. , 1 (N1), 1–34, (2000)

  7. [6]

    Aujol, Ch

    J.-F. Aujol, Ch. Dossal , Optimal rate of convergence of an ODE as- sociated to the fast gradient descent schemes for b > 0, Hal Preprint hal- 01547251, (2017)

  8. [7]

    Aujol, Ch

    J.-F. Aujol, Ch. Dossal, A. Rondepierre, Optimal convergence rates for Nesterov acceleration, SIAM J. Optim. , 29(4), 3131–3153, (2019)

Show all 45 references
  1. [8]

    Aujol, Ch

    J.-F. Aujol, Ch. Dossal, A. Rondepierre , Convergence rate of the heavy ball method for quasi-strongly convex optimization,SIAM J. Optim., 32(2), 1817–1842, (2022)

  2. [9]

    Arrow, A.C

    K.J. Arrow, A.C. Enthoven, Quasiconcave programming, Econometri- ca, 29, 779–800, (1961)

  3. [10]

    Generalized Con- cavity

    M. Avriel, W.E. Diewert, S. Schaible, I. Zang , “Generalized Con- cavity”. SIAM, Philadelphia, (2010)

  4. [11]

    Cabot, H

    A. Cabot, H. Engler, S. Gadat , On the long time behavior of second order differential equations with asymptotically small dissipation, Trans. Amer. Math. Soc. , 361, 5983–6017, (2009)

  5. [12]

    Generalized Convexity and Optimization: Theory and Applications

    A. Cambini, L. Martein . “Generalized Convexity and Optimization: Theory and Applications”. Springer, (2009)

  6. [13]

    Cherukuri, B

    A. Cherukuri, B. Gharesifard, J. Cort ´ es, Saddle-point dynamics: conditions for asymptotic stability of saddle points,SIAM J Control Optim , 55, 486–511, (2017)

  7. [14]

    Theory of value

    G. Debreu, “Theory of value”. John Wiley, New York, (1959)

  8. [15]

    Giselsson, S

    P. Giselsson, S. Boyd , Monotonicity and restart in fast gradient me- thods, In 53rd IEEE Conf. Decision and Control , 5058–5063, (2014)

  9. [16]

    Goudou, J

    X. Goudou, J. Munier , The gradient and heavy ball with friction dy- namical systems: the quasiconvex case, Math. Programm., 116, 173–191, (2009)

  10. [17]

    S.-M. Grad, F. Lara, R.T. Marcavillaca , Relaxed-inertial proximal point type algorithms for quasiconvex minimization, J. Global Optim. , 85, 615–635, (2023)

  11. [18]

    S.-M. Grad, F. Lara, R.T. Marcavillaca , Strongly quasiconvex functions: what we know (so far), J. Optim. Theory Appl. , DOI: 10.1007/s10957-025-02641-4, (2025). 26

  12. [19]

    Koumatos, S

    K. Koumatos, S. Spirito , Quasiconvex elastodynamics: weak-strong uniqueness for measure-valued solutions, Commun. Pure Appl. Math , 72, 1288–1320, (2019)

  13. [20]

    Handbook of Generali- zed Convexity and Generalized Monotonicity

    N. Hadjisavvas, S. Komlosi, S. Schaible , “Handbook of Generali- zed Convexity and Generalized Monotonicity”. Springer-Verlag, Bo ston, (2005)

  14. [21]

    Hermant, J.-F

    J. Hermant, J.-F. Aujol, C. Dossal, A. Rondepierre , Study of the behaviour of Nesterov accelerated gradient in a non convex settin g: the strongly quasar convex case. (2024). hal-04589853

  15. [22]

    Hinder, A

    O. Hinder, A. Sidford, N. Sohoni, Near-optimal methods for minimiz- ing star-convex functions and beyond. Proc. 33th Conf. on Lear. Theory , 125, 1894–1938, (2020)

  16. [23]

    Iusem, F

    A. Iusem, F. Lara, R. T. Marcavillaca, L.H. Yen, A two-step proxi- mal point algorithm for nonconvex equilibrium problems with application s to fractional programming. J. Global Optim. , 90, 3, 755–779, (2024)

  17. [24]

    Jovanovi ´c, Strongly quasiconvex quadratic functions, Publ

    M. Jovanovi ´c, Strongly quasiconvex quadratic functions, Publ. Inst. Math., Nouv. S´ er., 53, 153–156, (1993)

  18. [25]

    Jovanovi´c, A note on strongly convex and quasiconvex functions, Math

    M. Jovanovi´c, A note on strongly convex and quasiconvex functions, Math. Notes , 60, 584–585, (1996)

  19. [26]

    Machine Learning and Knowledge Discovery in Databases

    H. Karimi, J. Nutini, M. Schmidt , Linear convergence of gradient and proximal-gradient methods under the Polyak-/suppress Lojasiewicz condition. In “Machine Learning and Knowledge Discovery in Databases”, pages 7 95–

  20. [27]

    Korablev , Relaxation methods of minimization of pseudoconvex functions, J Soviet Math , 44, 1–5, (1989), (translated from Issled Prikl Mat, 8, 3–8, (1980))

    A.I. Korablev , Relaxation methods of minimization of pseudoconvex functions, J Soviet Math , 44, 1–5, (1989), (translated from Issled Prikl Mat, 8, 3–8, (1980))

  21. [28]

    Lara, On strongly quasiconvex functions: existence results and proxi- mal point algorithms, J

    F. Lara, On strongly quasiconvex functions: existence results and proxi- mal point algorithms, J. Optim. Theory Appl. , 192, 891–911, (2022)

  22. [30]

    F. Lara, C. Vega , Delayed feedback in online non-convex optimiza- tion: a non-stationary approach with applications. arXiv preprint, arXiv: 2412.14506, (2024)

  23. [31]

    /suppress Lojasiewicz, A topological property of real analytic subsets

    S. /suppress Lojasiewicz, A topological property of real analytic subsets. Coll. du CNRS, Les ´ equations aux d´ eriv´ ees partielles,117, 87–89, (1963). 27

  24. [32]

    S. N. Loizou, S. Vaswani, I.H. Laradji, S. Lacoste-Julien, Stochas- tic Polyak step-size for sgd: An adaptive learning rate for fast convergence, Int. Conf. Art. Intel. Stat. , PMLR, 1306–1314, (2021)

  25. [33]

    V. Mai, M. Johansson, Convergence of a stochastic gradient method with momentum for nonsmooth non-convex optimization, Int. Conf. Machine Lear., 6630–6639, (2020)

  26. [34]

    M.N. Nam, J. Sharkansky , On strong quasiconvexity of functions in infinite dimensions, arXiv: 2409.17450, (2024)

  27. [35]

    Nesterov, A method for solving the convex programming problem with convergence rate O (1 k2 ) , Dokl

    Y.E. Nesterov, A method for solving the convex programming problem with convergence rate O (1 k2 ) , Dokl. Akad. Nauk SSSR , 269, 543–547, (1983)

  28. [36]

    Introductory Lectures on Convex Optimization: A Basic Course

    Y.E. Nesterov, “Introductory Lectures on Convex Optimization: A Basic Course”. Springer, New York, (2004)

  29. [37]

    Nesterov, B.T

    Y.E. Nesterov, B.T. Polyak , Cubic regularization of newton method and its global performance, Math. Programm., 108(1), 177–205, (2006)

  30. [38]

    Piranfar, H

    M.R. Piranfar, H. Khatibzadeh , Long-Time Behavior of a Gradient System Governed by a Quasiconvex Function,J. Optim. Theory Appl., 188, 169–191, (2021)

  31. [39]

    Polyak, Gradient methods for minimizing functionals, Zh

    B.T. Polyak, Gradient methods for minimizing functionals, Zh. Vychisl. Math. Mat. Fiz. , 3, 643–653, (1963)

  32. [40]

    Polyak, Some methods of speeding up the convergence of iteration methods, USSR Comp

    B.T. Polyak, Some methods of speeding up the convergence of iteration methods, USSR Comp. Math. and Math. Phys. , 4(5), 1–17, (1964)

  33. [41]

    Polyak , Existence theorems and convergence of minimizing se- quences in extremum problems with restrictions, Soviet Math

    B.T. Polyak , Existence theorems and convergence of minimizing se- quences in extremum problems with restrictions, Soviet Math. , 7, 72–75, (1966)

  34. [42]

    P.W. Su, S. Boyd, E.J. Candes , A differential equation for modeling Nesterov’s accelerated gradient method: theory and insights, J. Mach. Learn. Res., 17(153), 1–43, (2016)

  35. [43]

    Inter- national conference on algorithmic learning theory. Proceedings of the 28th conference (ALT 2017)

    N. Thiemann, C. Igel, O. Wintenberger, Y. Seldin , A strongly quasiconvex PAC-Bayesian bound, in: Hanneke, Steve (ed.) et al., “ Inter- national conference on algorithmic learning theory. Proceedings of the 28th conference (ALT 2017)”, PMLR, 76, 466–492, (2017)

  36. [44]

    Vladimirov, Ju.E

    A.A. Vladimirov, Ju.E. Nesterov, Ju.N. Chekanov , O ravnomerno kvazivypuklyh funkcionalah [On uniformly quasiconvex functionals],Vestn. Mosk. un-ta, vycis. mat. i kibern. , 4, 18–27, (1978), (In Russian)

  37. [45]

    Vuong, A second-order dynamical system and its discretization for strongly pseudo-monotone variational inequality, SIAM J

    P.T. Vuong, A second-order dynamical system and its discretization for strongly pseudo-monotone variational inequality, SIAM J. Control Optim. , 59, 2875–2897, (2021). 28

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.