Pith. sign in

REVIEW 1 cited by

Algorithmic Regularization in Model-free Overparametrized Asymmetric Matrix Factorization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.02839 v2 pith:TVHFKFWZ submitted 2022-03-06 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords matrixapproximationdescentgradientinitializationobservedasymmetriccomplexity
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We study the asymmetric matrix factorization problem under a natural nonconvex formulation with arbitrary overparametrization. The model-free setting is considered, with minimal assumption on the rank or singular values of the observed matrix, where the global optima provably overfit. We show that vanilla gradient descent with small random initialization sequentially recovers the principal components of the observed matrix. Consequently, when equipped with proper early stopping, gradient descent produces the best low-rank approximation of the observed matrix without explicit regularization. We provide a sharp characterization of the relationship between the approximation error, iteration complexity, initialization size and stepsize. Our complexity bound is almost dimension-free and depends logarithmically on the approximation error, with significantly more lenient requirements on the stepsize and initialization compared to prior work. Our theoretical results provide accurate prediction for the behavior gradient descent, showing good agreement with numerical experiments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points

    cs.LG 2025-08 reject novelty 7.0 of 10

    Sub-n-gram estimators are near-stationary points of the population cross-entropy loss for in-context n-gram learning, offering a theoretical explanation for stage-wise training plateaus.

Pith tools