Pith. sign in

REVIEW 3 cited by

Stochastic Modified Equations and Dynamics of Dropout Algorithm

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.15850 v1 pith:65KAGURE submitted 2023-05-25 cs.LG

classification cs.LG
keywords dropoutstochasticequationsmodifieddynamicsflattermechanismminima
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Dropout is a widely utilized regularization technique in the training of neural networks, nevertheless, its underlying mechanism and its impact on achieving good generalization abilities remain poorly understood. In this work, we derive the stochastic modified equations for analyzing the dynamics of dropout, where its discrete iteration process is approximated by a class of stochastic differential equations. In order to investigate the underlying mechanism by which dropout facilitates the identification of flatter minima, we study the noise structure of the derived stochastic modified equation for dropout. By drawing upon the structural resemblance between the Hessian and covariance through several intuitive approximations, we empirically demonstrate the universal presence of the inverse variance-flatness relation and the Hessian-variance relation, throughout the training process of dropout. These theoretical and empirical findings make a substantial contribution to our understanding of the inherent tendency of dropout to locate flatter minima.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scalable Complexity Control Facilitates Reasoning Ability of LLMs

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Controlling model complexity through smaller initialization rates and stronger weight decay improved LLM benchmark scores and made loss-versus-scale curves descend faster.

  2. An Analysis for Reasoning Bias of Language Models with Small Initialization

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Initialization scale controls whether a transformer learns compositional reasoning or memorized mappings, because reasoning tokens acquire more differentiated embeddings early in training.

  3. Reasoning Bias of Next Token Prediction Training

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Training on all tokens (next token prediction) beats training only on answer tokens (critical token prediction) on small-scale reasoning benchmarks, an effect the authors attribute to noise-induced regularization.

Pith tools