REVIEW 6 cited by
When Are Nonconvex Problems Not Scary?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this note, we focus on smooth nonconvex optimization problems that obey: (1) all local minimizers are also global; and (2) around any saddle point or local maximizer, the objective has a negative directional curvature. Concrete applications such as dictionary learning, generalized phase retrieval, and orthogonal tensor decomposition are known to induce such structures. We describe a second-order trust-region algorithm that provably converges to a global minimizer efficiently, without special initializations. Finally we highlight alternatives, and open problems in this direction.
Forward citations
Cited by 6 Pith papers
-
Stochastic Saddle Avoidance Beyond Unit Excitation and Smoothness: A Pathwise Lyapunov-Perron Framework
A new pathwise Lyapunov-Perron framework proves almost sure saddle avoidance for stochastic recursions without unit excitation, covering SGD, mirror descent, proximal stochastic gradient, and random reshuffling.
-
Second-order methods for provably escaping strict saddle points in composite nonconvex and nonsmooth optimization
A trust-region method and a curvilinear linesearch method are shown to converge to second-order stationary points of composite nonconvex nonsmooth problems, independent of initialization.
-
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks
Gradient descent from moderately small random initialization jointly trains both layers of a ReLU network and converges linearly to the global minimizer for a linear target at order-wise optimal sample complexity by p...
-
Convergence of difference inclusions via a diameter criterion
A diameter criterion tied to a potential function certifies convergence of difference inclusions, enabling discrete proofs for first-order optimization methods with diminishing steps.
-
The Optimization Landscape of Carath\'eodory Decomposition of Toeplitz Covariances
Overparameterized gradient descent on a Carathéodory decomposition estimates Toeplitz covariances near the Cramér–Rao bound, and for fixed frequencies any stationary point of the amplitude objective recovers the true ...
-
Principles and Practice of Deep Representation Learning: or a Mathematical Theory of Memory
The book presents principles from optimization and information theory to explain deep network architectures and enable new interpretable models.
Discussion (0). Continue with ORCID to comment.