Pith. sign in

REVIEW 1 cited by

Eigen-componentwise convergence of SGD on quadratic programming

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.06476 v3 pith:U2BHKMZE submitted 2024-11-10 math.NA cs.NA

classification math.NAcs.NA
keywords convergencecomponentssingularvaluescorrespondingerrorfasterinitial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Stochastic gradient descent (SGD) is a workhorse algorithm for solving large-scale optimization problems in data science and machine learning. Understanding the convergence of SGD is hence of fundamental importance. In this work we examine the SGD convergence (with various step sizes) when applied to unconstrained convex quadratic programming (essentially least-squares (LS) problems), and in particular analyze the error components respect to the eigenvectors of the Hessian. The main message is that the convergence depends largely on the corresponding eigenvalues (singular values of the coefficient matrix in the LS context), namely the components for the large singular values converge faster in the initial phase. We then show there is a phase transition in the convergence where the convergence speed of the components, especially those corresponding to the larger singular values, will decrease. Finally, we show that the convergence of the overall error (in the solution) tends to decay as more iterations are run, that is, the initial convergence is faster than the asymptote.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hard edge asymptotics of correlation functions between singular values and eigenvalues

    math.PR 2025-01 conditional novelty 7.0 of 10

    For a broad class of bi-unitarily invariant random matrix ensembles, the large-n limit of the joint density of one eigenradius and k singular values at the hard edge is expressed through the limiting kernel of the sin...

Pith tools