Pith. sign in

REVIEW 2 cited by

Almost Sure Convergence of Dropout Algorithms for Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.02247 v2 pith:V62ZN5TO submitted 2020-02-06 math.OC cs.LGmath.PR

classification math.OCcs.LGmath.PR
keywords dropoutconvergenceratealgorithmsdepthfunctionsprobabilityactivation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We investigate the convergence and convergence rate of stochastic training algorithms for Neural Networks (NNs) that have been inspired by Dropout (Hinton et al., 2012). With the goal of avoiding overfitting during training of NNs, dropout algorithms consist in practice of multiplying the weight matrices of a NN componentwise by independently drawn random matrices with $\{0, 1 \}$-valued entries during each iteration of Stochastic Gradient Descent (SGD). This paper presents a probability theoretical proof that for fully-connected NNs with differentiable, polynomially bounded activation functions, if we project the weights onto a compact set when using a dropout algorithm, then the weights of the NN converge to a unique stationary point of a projected system of Ordinary Differential Equations (ODEs). After this general convergence guarantee, we go on to investigate the convergence rate of dropout. Firstly, we obtain generic sample complexity bounds for finding $\epsilon$-stationary points of smooth nonconvex functions using SGD with dropout that explicitly depend on the dropout probability. Secondly, we obtain an upper bound on the rate of convergence of Gradient Descent (GD) on the limiting ODEs of dropout algorithms for NNs with the shape of arborescences of arbitrary depth and with linear activation functions. The latter bound shows that for an algorithm such as Dropout or Dropconnect (Wan et al., 2013), the convergence rate can be impaired exponentially by the depth of the arborescence. In contrast, we experimentally observe no such dependence for wide NNs with just a few dropout layers. We also provide a heuristic argument for this observation. Our results suggest that there is a change of scale of the effect of the dropout probability in the convergence rate that depends on the relative size of the width of the NN compared to its depth.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dropout Neural Network Training Viewed from a Percolation Perspective

    cs.LG 2025-12 conditional novelty 6.0 of 10

    Dropout can break training in deep narrow no-bias networks because the probability of a surviving input–output path decays so fast that gradients vanish almost always.

  2. In situ fine-tuning of in silico trained Optical Neural Networks

    cs.NE 2025-06 reject novelty 6.0 of 10

    GIFT computes a gradient-informed direction based on how the training loss gradient changes with the assumed noise level and line-searches along it in situ, improving accuracy under noise misspecification.

Pith tools