REVIEW 1 cited by
The Global Landscape of Neural Networks: An Overview
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
One of the major concerns for neural network training is that the non-convexity of the associated loss functions may cause bad landscape. The recent success of neural networks suggests that their loss landscape is not too bad, but what specific results do we know about the landscape? In this article, we review recent findings and results on the global landscape of neural networks. First, we point out that wide neural nets may have sub-optimal local minima under certain assumptions. Second, we discuss a few rigorous results on the geometric properties of wide networks such as "no bad basin", and some modifications that eliminate sub-optimal local minima and/or decreasing paths to infinity. Third, we discuss visualization and empirical explorations of the landscape for practical neural nets. Finally, we briefly discuss some convergence results and their relation to landscape results.
Forward citations
Cited by 1 Pith paper
-
Noise-Driven Exploration and Transient Freezing Select Flat Minima in Stochastic Gradient Descent
SGD's preference for flat minima is explained by a noise-controlled transient exploration phase that ends in a freezing transition; stronger noise delays freezing and biases selection toward flatter valleys.
Discussion (0). Continue with ORCID to comment.