Pith. sign in

REVIEW 2 cited by

SGD on Neural Networks Learns Functions of Increasing Complexity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.11604 v1 pith:26CAGDVH submitted 2019-05-28 cs.LG cs.NEstat.ML

classification cs.LGcs.NEstat.ML
keywords classifiercomplexityevenfunctionshypothesisincreasinginitiallearns
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We perform an experimental study of the dynamics of Stochastic Gradient Descent (SGD) in learning deep neural networks for several real and synthetic classification tasks. We show that in the initial epochs, almost all of the performance improvement of the classifier obtained by SGD can be explained by a linear classifier. More generally, we give evidence for the hypothesis that, as iterations progress, SGD learns functions of increasing complexity. This hypothesis can be helpful in explaining why SGD-learned classifiers tend to generalize well even in the over-parameterized regime. We also show that the linear classifier learned in the initial stages is "retained" throughout the execution even if training is continued to the point of zero training error, and complement this with a theoretical result in a simplified model. Key to our work is a new measure of how well one classifier explains the performance of another, based on conditional mutual information.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Conditioning LLM training on natural-language evaluator descriptions and deploying with a held-out higher-fidelity label improved even-handedness and reduced sycophancy in two small proof-of-concept experiments.

  2. Mitigating Spurious Correlations with Memorization-Guided Dataset De-Biasing

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Proposes memorization-guided two-stage scoring to select debiased training subsets, enabling ERM models to achieve better performance than SOTA debiasing techniques using only 10% of data.

Pith tools