Pith. sign in

REVIEW 1 cited by

Implicit Regularization in Deep Learning May Not Be Explainable by Norms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.06398 v2 pith:QF4LYOTX submitted 2020-05-13 cs.LG cs.NEstat.ML

classification cs.LGcs.NEstat.ML
keywords implicitnormsregularizationmatrixdeepfactorizationlearninginterpretation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Mathematically characterizing the implicit regularization induced by gradient-based optimization is a longstanding pursuit in the theory of deep learning. A widespread hope is that a characterization based on minimization of norms may apply, and a standard test-bed for studying this prospect is matrix factorization (matrix completion via linear neural networks). It is an open question whether norms can explain the implicit regularization in matrix factorization. The current paper resolves this open question in the negative, by proving that there exist natural matrix factorization problems on which the implicit regularization drives all norms (and quasi-norms) towards infinity. Our results suggest that, rather than perceiving the implicit regularization via norms, a potentially more useful interpretation is minimization of rank. We demonstrate empirically that this interpretation extends to a certain class of non-linear neural networks, and hypothesize that it may be key to explaining generalization in deep learning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Grokking Beyond the Euclidean Norm of Model Parameters

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Grokking is induced by any small nonzero regularizer whose favored solutions generalize, with a delay that scales like one over the learning rate times the regularization strength.

Pith tools