Pith. sign in

REVIEW 2 cited by

A convergence result of a continuous model of deep learning via a \L{}ojasiewicz--Simon inequality

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.15365 v3 pith:VCTXVKTH submitted 2023-11-26 cs.LG math.APmath.FAmath.PR

classification cs.LGmath.APmath.FAmath.PR
keywords objectiveinequalitytraininganalyticitycontinuousconvergencecurvedeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study an idealized training process for deep neural networks in a continuous-depth, mean-field model in which each layer is parameterized by a probability measure on a Euclidean parameter space. The training dynamics are formulated as a Wasserstein-type gradient flow of an objective with a fixed $L^2$-regularization. Under suitable analyticity and growth assumptions, together with a coercivity assumption and sufficient regularity of the initial data, we prove that every curve of maximal slope converges to a single critical point of the objective as the training time tends to infinity. The proof combines compactness of the curve with a \L{}ojasiewicz--Simon inequality for the metric slope. To establish the inequality, we lift the objective to a Hilbert space of random variables and use the analyticity of the lifted gradient in a stronger $L^\infty$ topology to overcome its lack of continuous differentiability in the Hilbert-space topology. Our convergence result does not require global displacement convexity, a Polyak--\L{}ojasiewicz-type condition, or initialization near a minimizer; the objective may remain genuinely nonconvex.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Genericity of Polyak-Lojasiewicz Inequalities for Entropic Mean-Field Neural ODEs

    math.OC 2025-07 conditional novelty 7.0 of 10

    For a continuum-depth ResNet model with entropic regularization, an open dense set of initial feature-label distributions admits a unique stable minimizer and a local Polyak-Lojasiewicz inequality near it.

  2. Dropout and Random Gradient Masking Are Asymptotically Equivalent in Large ResNets

    stat.ML 2026-07 conditional novelty 6.0 of 10

    As ResNets grow deep and wide with fixed dropout rate, dropout training and random-gradient-masking training converge to the same limiting dynamics, and the common masking variants collapse to one limit.

Pith tools