Pith. sign in

REVIEW 2 cited by

On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.12470 v2 pith:ENNUOSLP submitted 2021-02-24 cs.LG stat.ML

classification cs.LGstat.ML
keywords approximationdeepdifferentialequationsgeneralizationnetssdessimulation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

It is generally recognized that finite learning rate (LR), in contrast to infinitesimal LR, is important for good generalization in real-life deep nets. Most attempted explanations propose approximating finite-LR SGD with Ito Stochastic Differential Equations (SDEs), but formal justification for this approximation (e.g., (Li et al., 2019)) only applies to SGD with tiny LR. Experimental verification of the approximation appears computationally infeasible. The current paper clarifies the picture with the following contributions: (a) An efficient simulation algorithm SVAG that provably converges to the conventionally used Ito SDE approximation. (b) A theoretically motivated testable necessary condition for the SDE approximation and its most famous implication, the linear scaling rule (Goyal et al., 2017), to hold. (c) Experiments using this simulation to demonstrate that the previously proposed SDE approximation can meaningfully capture the training and generalization properties of common deep nets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Superlinear Relationship between SGD Noise Covariance and Loss Landscape Curvature

    cs.LG 2026-02 reject novelty 5.0 of 10

    SGD noise covariance is claimed to follow the second moment of per-sample Hessians, giving a superlinear power law C_ii ∝ H_ii^γ with 1 ≤ γ ≤ 2.

  2. Limit Theorems for Stochastic Gradient Descent in High-Dimensional Single-Layer Networks

    stat.ML 2025-11 unverdicted novelty 5.0 of 10

    At the critical step-size scaling for SGD in high-dimensional single-layer networks, effective dynamics gain a diffusive correction term that changes the phase diagram and reduces to an Ornstein-Uhlenbeck process near...

Pith tools